Design an online bookstore: DB insert succeeds but SQS publish fails, and duplicate messages
Company: Databricks
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design the backend of an online bookstore where customers browse books and place orders. When an order request arrives, the service inserts the order into its database and then publishes a message to a queue (Amazon SQS) so that other parts of the system can process the order asynchronously.
The overall flow is expected to be covered quickly. Most of the interview goes into failure modes: what should happen when the database insert succeeds but the publish to SQS fails, how to prevent that situation, and how to keep consumers from processing the same message twice when messages are delivered at least once.
### Constraints and Clarifications
- Assume the order store is a database with transactions (for example, a relational database) unless told otherwise. The queue is Amazon SQS, and messages are delivered at least once.
- Scale, the full feature list and the set of downstream consumers are not given. Ask for them or state your assumptions.
- Expect to be challenged on each mechanism you propose. Describe how it works step by step; a pattern's name alone is not an answer.
### Clarifying Questions
- Which features are in scope: catalog browsing and search, cart, checkout and payment, inventory, order history, notifications?
- What consumes the order messages, and what does a lost or a duplicated message cost for each consumer?
- Must the customer see the order confirmed immediately, or is "order received, processing" acceptable?
- What order volume should the design handle, and what does a peak (for example, a popular new release) look like?
- Is per-order ordering of messages required, and is a standard or a FIFO queue expected?
### Part 1 — Core design and order flow
Design the main components, the data model and the APIs, and walk through the path of an order from the customer's request to the asynchronous processing behind the queue.
```hint Where the order becomes real
Decide which single write makes an order exist, and what every later step may assume once that write has happened.
```
```hint The last copy
Two customers buy the last copy of a book at the same moment. Decide where that conflict is resolved.
```
#### What This Part Should Cover
- Components, data model and APIs for the catalog, orders and inventory
- The synchronous request path versus the work done behind the queue
- Order states and how the customer sees progress
- Correct stock handling when orders for the same book arrive concurrently
### Part 2 — The database insert succeeds, the SQS publish fails
A request inserts an order into the database, but the following publish to SQS fails. What should happen, and how can you prevent the database and the queue from disagreeing?
```hint Two writes, no shared transaction
The database and SQS cannot commit together. List every point where a call can fail or the process can crash between the two writes, and the state each one leaves behind.
```
```hint Survive the crash
Retrying in the request handler only helps if the process stays alive. Think about what durable record would let another process finish the job later.
```
#### Clarifying Questions for this Part
- If the message cannot be sent, may the order be rolled back, or must it be kept?
- How much delay is acceptable between committing an order and downstream services seeing it?
#### What This Part Should Cover
- The failure cases: the publish fails, the process crashes between the two writes, and the order of the two writes is reversed
- A mechanism that guarantees every committed order is eventually published, explained step by step
- Alternatives and their trade-offs
- Operating the chosen mechanism: delay, cleanup and monitoring
### Part 3 — Duplicate messages under at-least-once delivery
Messages are delivered at least once. How do the consumers avoid processing the same order message twice?
```hint Where duplicates come from
Count the ways a message can arrive twice: from your own publishing side, and from the queue itself.
```
```hint Make the second time harmless
For each consumer, decide what key identifies "this work was already done", and where and when that fact is recorded.
```
#### What This Part Should Cover
- The sources of duplicate messages, on the publishing side and in the queue
- A stable message key, and how a consumer records processed keys together with its effect
- Side effects outside the consumer's own database, such as payments and emails
- What queue-level deduplication can and cannot guarantee
### What a Strong Answer Covers
- A clear order flow whose single source of truth is the database commit
- No lost messages: every committed order reaches the queue eventually, even across crashes
- Idempotent consumers, so at-least-once delivery produces each effect once
- Mechanisms explained concretely enough that a listener who does not know the pattern names can follow them
- Retries, dead-letter handling, monitoring and reconciliation
### Follow-up Questions
- The process that forwards messages to SQS crashes after SQS accepts a message but before it records that the message was sent. What happens next, and why is that acceptable?
- A consumer's processing takes longer than the SQS visibility timeout. What goes wrong, and how do you fix it?
- Why not publish to SQS first and insert into the database second?
- How would you detect an order that was committed but never processed downstream?
Overview: System design question about an online bookstore whose order service inserts orders into a database and then publishes messages to Amazon SQS. Probes what should happen when the insert succeeds but the publish fails, how to keep the database and queue consistent across crashes, and how consumers avoid duplicate processing under at-least-once delivery.
Read the full Databricks Software Engineer interview experience this question came from