Design a Scalable Error Message Counting Service

Quick Overview

Design a service that receives application error messages and reports how often each distinct error occurs over time. It tests grouping varied messages into error signatures, absorbing bursty write traffic during incidents, time-windowed aggregation, top-error queries and count accuracy under failures.

Design a Scalable Error Message Counting Service

Company: LinkedIn

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design an error message counter: a service that receives the error messages applications emit and reports how many times each error has occurred. Treat it as a shared service that many applications report to, and expect the same underlying error to produce messages that differ in details such as request IDs, user IDs or timestamps. ```hint Decide what one error is Two messages from the same failure rarely match byte for byte. Settle what key you count under before you decide where the counts are stored. ``` ```hint Plan for the incident, not the average Error traffic is bursty: one bad deploy can multiply it within seconds, and that is exactly when people read the counts. Think about where the write path absorbs a burst. ``` ### Clarifying Questions - Which queries must be supported: the count for one error over a time range, the most frequent errors in a recent window, a time series per error, or all of these? - Are counts broken down by service, host, region or release version? - How many applications report errors, and what are the normal and peak error rates? - How fresh must the counts be, and must they be exact, or is a small bounded error acceptable? - How long must counts be kept, and at what time granularity? - Is the counter also expected to raise alerts, or only to answer queries? ### What a Strong Answer Covers - A precise counting key: how messages are normalized into error signatures, and which dimensions and time buckets are counted. - A write path that survives bursts, with aggregation before storage rather than one database write per error. - A storage layout and query path for time-range counts and top errors, including rollups and retention. - Correctness under retries, late arrivals and component failures, with an explicit accuracy guarantee. - Back-of-the-envelope estimates that justify the choices, and the signals used to operate the system. ### Follow-up Questions - How would you alert when a new error signature appears, or when a known one rises well above its usual rate? - A single error suddenly accounts for most of the traffic. Which component becomes a hot spot, and how do you spread the load? - If normalization misses a variable part of a message and the number of distinct signatures explodes, what breaks and how do you contain it? - How would you let a team merge two signatures that turn out to be the same error, without losing historical counts?

Overview: Design a service that receives application error messages and reports how often each distinct error occurs over time. It tests grouping varied messages into error signatures, absorbing bursty write traffic during incidents, time-windowed aggregation, top-error queries and count accuracy under failures.

|Home/System Design/LinkedIn
LinkedIn logo
LinkedIn
Sep 24, 2026
mediumSoftware EngineerOnsiteSystem Design
1
0

Design an error message counter: a service that receives the error messages applications emit and reports how many times each error has occurred. Treat it as a shared service that many applications report to, and expect the same underlying error to produce messages that differ in details such as request IDs, user IDs or timestamps.

Clarifying Questions Guidance

  • Which queries must be supported: the count for one error over a time range, the most frequent errors in a recent window, a time series per error, or all of these?
  • Are counts broken down by service, host, region or release version?
  • How many applications report errors, and what are the normal and peak error rates?
  • How fresh must the counts be, and must they be exact, or is a small bounded error acceptable?
  • How long must counts be kept, and at what time granularity?
  • Is the counter also expected to raise alerts, or only to answer queries?

What a Strong Answer Covers Guidance

  • A precise counting key: how messages are normalized into error signatures, and which dimensions and time buckets are counted.
  • A write path that survives bursts, with aggregation before storage rather than one database write per error.
  • A storage layout and query path for time-range counts and top errors, including rollups and retention.
  • Correctness under retries, late arrivals and component failures, with an explicit accuracy guarantee.
  • Back-of-the-envelope estimates that justify the choices, and the signals used to operate the system.

Follow-up Questions Guidance

  • How would you alert when a new error signature appears, or when a known one rises well above its usual rate?
  • A single error suddenly accounts for most of the traffic. Which component becomes a hot spot, and how do you spread the load?
  • If normalization misses a variable part of a message and the number of distinct signatures explodes, what breaks and how do you contain it?
  • How would you let a team merge two signatures that turn out to be the same error, without losing historical counts?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...