Design a Simple Message Queue Service Without an Off-the-Shelf Broker
Company: Attentive
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a simple message queue service using simpler primitives. You may not build it on an off-the-shelf queue or log such as Kafka: the point is to explain the underlying mechanisms directly, rather than delegating them to an existing product.
The setup:
- A set of producers sends messages to the queue.
- The queue is a service, not a database. It receives messages from the producers, and you define the interface between producers and the service: HTTP, WebSocket, or anything else.
- The queue stores the messages in some order. The content of a message does not matter for this problem.
- Consumers on the other side also connect to the queue service, and you define that interface as well. For example, once four messages have been published, connected consumers read them from the service.
Expect requirements to be added or changed as the design progresses. Part of the exercise is to re-check earlier decisions each time that happens.
### Clarifying Questions
- Should each message go to exactly one consumer (a work queue), to every consumer (publish-subscribe), or should both be possible through consumer groups?
- What ordering is required: one global order, order per producer, or order per key?
- Which delivery guarantee is required: at most once, at least once, or effectively once?
- Must an accepted message survive a server crash, and how long are messages retained after they are consumed?
- The problem was introduced under the heading of a metrics system. Are the messages metrics events, and does that relax ordering or loss tolerance?
- What message rates, message sizes, and numbers of producers and consumers should the design support?
### Part 1 — Producer interface and ordered storage
Define the producer-facing interface and how the service accepts, orders, and stores messages. Explain exactly what a producer can rely on once the service acknowledges a message.
```hint What does an acknowledgment promise?
Decide the exact moment the service is allowed to say a message is accepted, and what has to be true on disk at that moment.
```
#### What This Part Should Cover
- The producer API, including batching, retries, and backpressure
- How order is assigned and preserved in storage
- Durability: when the acknowledgment is sent, and what survives a crash
### Part 2 — Consumer interface and delivery
Define how consumers connect and receive messages, how the service learns that a message has been processed, and what happens when a consumer fails in the middle of a message.
```hint Track progress, not just messages
Think about the minimal per-consumer state that lets the service resume correctly after either a consumer or the service itself restarts.
```
#### What This Part Should Cover
- Push versus pull, and flow control for slow consumers
- Acknowledgment, redelivery, and handling of a message that always fails
- How the delivery and ordering guarantees agreed in the clarifications are actually met
### Part 3 — Troubleshooting
Once the queue is running, walk through how you would detect and diagnose problems such as consumers falling behind, messages reported missing, or messages processed twice.
```hint Make every message traceable
Ask which identifiers and counters would let you follow one message from the producer's request to the consumer's acknowledgment.
```
#### What This Part Should Cover
- Metrics and logs that make lag, loss, and duplication visible
- A systematic diagnosis path for each symptom
- Which design choices from Parts 1 and 2 cause each symptom, and the fix
### What a Strong Answer Covers
- Ordering, delivery, and durability guarantees stated up front and re-checked whenever a requirement changes
- Mechanisms built from basic primitives such as files, offsets, and acknowledgments, rather than named products
- Explicit trade-offs: latency versus durability, ordering versus parallelism, push versus pull
- Failure handling for producer retries, consumer crashes, and service restarts
- The candidate driving the technical discussion and covering small operational details, not only the high-level boxes
### Follow-up Questions
- How would you scale beyond one server while keeping the ordering guarantee you promised?
- How would you replicate the queue so that losing one server loses no acknowledged message?
- What happens to disk usage and retention when one consumer group stops consuming for a day?
Overview: Design a simple message queue service from basic primitives rather than an off-the-shelf broker, defining the producer and consumer interfaces yourself. It tests ordered durable storage, acknowledgment and redelivery semantics, adapting to shifting requirements, and troubleshooting lag, loss and duplicate delivery.
Read the full Attentive Software Engineer interview experience this question came from