Design Durable One-to-One Web Messaging at Large Scale
Company: Anthropic
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Technical Screen
Design a one-to-one text messaging application with Web clients and 100 million daily active users. Group chats, images, video, and mobile push notifications are out of scope. Messages for an offline recipient must be retained and available when that recipient reconnects.
### Constraints & Assumptions
- Persist messages, define per-conversation ordering, handle retries, and support low-latency online delivery.
- The report uses **100 sent messages per active user per day as a sizing assumption**, not as a fixed product requirement. Analyze its consequences and identify what additional measurements are needed.
- Daily active users do not specify peak simultaneous WebSocket connections, message size, retention, or traffic skew.
- Distinguish server acceptance, durable storage, client receipt, and human-read status.
### Clarifying Questions to Ask
- What acknowledgment means success to the sender, and what durability does it promise?
- Is ordering required within each chat, and are gaps in assigned sequence numbers possible?
- What message retention, average payload size, and peak factor should storage planning use?
- Can a user have several active browser connections, and is last-received state distinct from last-read state?
### Part 1 — Size the Durable Message Path
Estimate writes under the stated assumption and design the message store and shard key. Explain which kinds of load adding app servers, caches, or read replicas actually relieves.
#### What This Part Should Cover
- Daily and average per-second writes, peak uncertainty, and write amplification.
- Chat-scoped partitioning and ordering.
- Separation of connection capacity, durable writes, history reads, and presence traffic.
### Part 2 — Deliver One Online or Offline Message
Trace a message from the sender's WebSocket through durable storage, sender acknowledgment, presence lookup, forwarding, recipient delivery, and reconnect recovery.
#### What This Part Should Cover
- A stable client operation identity and durable deduplication across server restarts.
- Per-chat sequence and a recoverable delivery intent or catch-up mechanism.
- Gaps, duplicates, stale presence, and forwarding failures.
### Part 3 — Find a User's Chats and Missing Messages
Specify the user-to-chat mapping, indexes, and cursor queries. Explain read-replica lag for historical reads versus immediate online gap repair.
#### What This Part Should Cover
- A concrete keyed membership query rather than scanning all chats.
- Separate list metadata and authoritative message history.
- Replica freshness or primary fallback where a recently committed message is needed.
### Part 4 — Handle a Traffic Burst
Prioritize work and identify which bottleneck each scaling or degradation measure addresses.
#### What This Part Should Cover
- Durable-send admission before unbounded queues form.
- Bounded online delivery and reconnect load.
- Historical-read degradation and caches without claiming they eliminate message writes.
```hint A server-local duplicate buffer can disappear
A retried send may reach another app server after the original server restarts. Decide where the idempotency decision survives that change.
```
### What a Strong Answer Covers
- A complete durable lifecycle for online and offline messages.
- Capacity and storage decisions grounded in the explicitly labeled traffic assumption.
- Concrete membership and catch-up queries, correct deduplication, and replica-aware recovery.
### Follow-up Questions
- How should the client recover if the live channel drops the final message and no later sequence exposes the gap?
- What extra writes arise if every message immediately updates both users' chat-list summaries?
- How can multiple browser tabs avoid confusing “received on one device” with “read by the user”?
Overview: Design Web-only one-to-one messaging at 100 million DAU with durable deduplication, chat sharding, membership indexes, offline recovery, and replica-aware reads.
Read the full Anthropic Software Engineer interview experience this question came from