Design a Scalable Error Message Counting Service
Company: LinkedIn
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design an error message counter: a service that receives the error messages applications emit and reports how many times each error has occurred. Treat it as a shared service that many applications report to, and expect the same underlying error to produce messages that differ in details such as request IDs, user IDs or timestamps.
```hint Decide what one error is
Two messages from the same failure rarely match byte for byte. Settle what key you count under before you decide where the counts are stored.
```
```hint Plan for the incident, not the average
Error traffic is bursty: one bad deploy can multiply it within seconds, and that is exactly when people read the counts. Think about where the write path absorbs a burst.
```
### Clarifying Questions
- Which queries must be supported: the count for one error over a time range, the most frequent errors in a recent window, a time series per error, or all of these?
- Are counts broken down by service, host, region or release version?
- How many applications report errors, and what are the normal and peak error rates?
- How fresh must the counts be, and must they be exact, or is a small bounded error acceptable?
- How long must counts be kept, and at what time granularity?
- Is the counter also expected to raise alerts, or only to answer queries?
### What a Strong Answer Covers
- A precise counting key: how messages are normalized into error signatures, and which dimensions and time buckets are counted.
- A write path that survives bursts, with aggregation before storage rather than one database write per error.
- A storage layout and query path for time-range counts and top errors, including rollups and retention.
- Correctness under retries, late arrivals and component failures, with an explicit accuracy guarantee.
- Back-of-the-envelope estimates that justify the choices, and the signals used to operate the system.
### Follow-up Questions
- How would you alert when a new error signature appears, or when a known one rises well above its usual rate?
- A single error suddenly accounts for most of the traffic. Which component becomes a hot spot, and how do you spread the load?
- If normalization misses a variable part of a message and the number of distinct signatures explodes, what breaks and how do you contain it?
- How would you let a team merge two signatures that turn out to be the same error, without losing historical counts?
Overview: Design a service that receives application error messages and reports how often each distinct error occurs over time. It tests grouping varied messages into error signatures, absorbing bursty write traffic during incidents, time-windowed aggregation, top-error queries and count accuracy under failures.