Design a Notification System for Bursty Writes of 1M per Second
Company: Coreweave
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a notification system. The interviewer pasted a long list of requirements that is not available here, so confirm the requirements below or state your own.
When you analyze scale and traffic, assume the interviewer's figure: about 1 million writes per second, arriving in bursts. The system must therefore absorb bursty, write-heavy load. Expect to go through the functional and non-functional requirements end to end, to justify your storage choice (in this round: why a store such as Cassandra writes quickly), and to handle duplicate notifications.
### Constraints and Clarifications
- Peak write traffic is about 1 million writes per second, and it is bursty.
- Average traffic, retention, channels and delivery guarantees are not given; ask for them or state your assumptions.
### Clarifying Questions
- What counts as a write: a notification request from an upstream service, a per-recipient notification record, or a delivery-status update?
- Which channels are in scope: mobile push, email, SMS, an in-app inbox?
- What delivery semantics are required: at least once with deduplication, or at most once? How fast must a notification arrive, and does that differ by priority?
- Do users have preferences, quiet hours or rate limits? Are there priority levels, such as security alerts versus marketing?
- Must users be able to read past notifications, and for how long?
- How bursty is the traffic: what is the ratio of peak to average, and how long does a burst last?
### Part 1 — Requirements, scale and traffic pattern
State the functional and non-functional requirements, estimate the load from the 1 million writes per second peak, describe the traffic pattern, and derive what each part of the architecture must absorb.
```hint Peak versus average
Separate the components that must accept a write at the peak rate from those that can drain a backlog more slowly after the burst.
```
#### What This Part Should Cover
- Functional requirements specific to notifications: channels, preferences, inbox, delivery status
- Non-functional requirements: latency by priority, durability, availability, deduplication
- Back-of-the-envelope throughput, bandwidth and storage
- Which components are sized for peak and which for average
### Part 2 — Absorbing bursty heavy writes
Design the ingestion and storage path. If your design uses a wide-column store such as Cassandra, explain why its writes are fast and what that costs on reads.
```hint Accept now, process later
Decide what must happen before the API acknowledges a write, and what can safely be deferred.
```
```hint How writes reach disk
Ask how the store turns a stream of writes into disk operations, and which work it postpones until later.
```
#### What This Part Should Cover
- A durable, partitioned buffer between the API and processing, with partition keys and backpressure
- The storage data model and partition keys that avoid hot partitions
- The mechanics of a log-structured write path and its read and compaction costs
- Sizing of the buffer and storage clusters
### Part 3 — Delivery and deduplication
Design delivery to the channels, and make sure a user does not receive the same notification twice when requests or retries are duplicated.
```hint Where duplicates come from
List every step that may retry, from the upstream caller to the channel provider.
```
```hint Remembering is not free
Decide which key identifies a duplicate and how long it must be remembered at this write rate.
```
#### What This Part Should Cover
- Idempotency keys and a deduplication window held in a cache, with its memory cost
- At-least-once processing with idempotent consumers and delivery-state tracking
- Channel workers, provider rate limits, retries with backoff and a dead-letter queue
- Preferences and priority isolation during bursts
### What a Strong Answer Covers
- Requirements and estimates that drive concrete design choices
- Burst absorption without dropping writes or slowing acknowledgments
- A storage choice justified by its write-path mechanics, with its costs stated
- Delivery semantics stated and achieved, including deduplication
- Failure handling and observability across the pipeline
### Follow-up Questions
- A burst at 1 million writes per second lasts ten minutes, and the push provider accepts far less. What happens to the backlog, and how do urgent notifications get through?
- The deduplication cache loses its data during a failover. Which duplicates can reach users, and how do you limit them?
- Why not store notifications in PostgreSQL? At what point does it stop working for this load?
- How would you send one notification to tens of millions of users at once?
Overview: System design question for a notification system that must absorb bursty traffic of about one million writes per second. It covers requirements and capacity estimates, a queue-buffered write path, why a log-structured store such as Cassandra writes quickly, channel delivery and deduplication.
Read the full Coreweave Software Engineer interview experience this question came from