Design scalable notification system
Company: Airbnb
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
##### Question
Design a scalable notification system for a marketplace at Airbnb scale. It must deliver messages to millions of users across **email, SMS, and push (APNs/FCM)** — with an optional social/third-party channel — and serve several recipient types (guests, hosts, and internal employees).
The system supports two delivery modes:
- **Event-triggered (single) notifications** — fired by a real-time domain event, e.g. `BookingConfirmed` sends an email to the guest, or `CheckInDateReached` sends the door-lock code by SMS.
- **Bulk / batch campaigns** — a scheduled or ad-hoc send to a target audience, e.g. a promotion to potential hosts.
Assume existing product services already emit domain events (booking created, check-in reached, campaign created); your system consumes them and delivers the notifications.
Walk through an end-to-end design covering:
1. **Requirements gathering and scale assumptions** — functional and non-functional requirements, throughput and latency targets, reliability guarantees.
2. **High-level architecture and data flow** — core components, and the end-to-end path for *both* the event-triggered and the batch/campaign modes.
3. **Data model** — the key entities and their relationships.
4. **API design** — the producer/ingestion API, campaign API, and provider webhooks.
5. **Templates and personalization** — how messages are authored, versioned, localized, rendered, and routed to a channel.
6. **User preferences and compliance** — per-channel opt-in/out, quiet hours, frequency caps, unsubscribe/STOP handling, suppression lists.
7. **Message prioritization** — keeping transactional traffic (e.g. an OTP or a lock code) from being starved by marketing campaigns.
8. **Deduplication, idempotency, and ordering** — preventing double sends, and enforcing per-user ordering where it actually matters.
9. **Retries and failure handling** — retry policy by error class, dead-letter queues, provider outages and failover.
10. **Scaling strategies** — partitioning, fan-out for large campaigns, rate limiting, backpressure, multi-region.
11. **Observability** — metrics/SLIs, tracing, logging, alerting, and the compliance audit trail ("why did this user get this message?").
12. **Cost considerations** — per-channel unit economics and the controls that keep spend predictable.
Be explicit about your assumptions, the trade-offs behind each choice, and how you would validate the design before it carries production traffic.
Overview: Airbnb's onsite system design round asks you to design a notification system that delivers email, SMS, and push messages to millions of guests, hosts, and employees, supporting both event-triggered sends and large batch campaigns. A complete answer covers requirements and scale math, architecture and data flow for both modes, the data model and API, templating, user preferences and compliance, prioritization, deduplication and idempotency, retries and failover, scaling and backpressure, observability, and cost. This walkthrough gives a corrected end-to-end reference answer for every part, including the trade-offs interviewers probe.