Design scalable notification system

Quick Overview

Design scalable notification system evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

Design scalable notification system

Company: Airbnb

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

##### Question Design a scalable notification system that can send messages (email, SMS, push) to millions of users with low latency. Cover requirements gathering, high-level architecture, data model, API design, message prioritization, deduplication, retries, failure handling, scaling strategies, monitoring, and cost considerations.

Quick Answer: Design scalable notification system evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

|Home/System Design/Airbnb
Airbnb logo
Airbnb
Jul 29, 2025, 8:05 AM
hardSoftware EngineerOnsiteSystem Design
31
0

Design scalable notification system

System Design: Low-Latency, Multi-Channel Notification Platform

You are asked to design a scalable, reliable notification system that can send messages to millions of users with low latency across multiple channels (email, SMS, push).

Assume a large, consumer-facing product operating globally with both transactional (real-time) and bulk/marketing use cases.

Cover the following:

  1. Requirements Gathering
    • Functional, non-functional, traffic assumptions, SLAs/latency targets, compliance.
  2. High-Level Architecture
    • Core components and data flow for real-time, scheduled, and bulk sends.
  3. Data Model
    • Key entities (templates, preferences, messages, attempts, providers, etc.).
  4. API Design
    • Producer APIs, admin APIs, idempotency, status callbacks/webhooks.
  5. Message Prioritization
    • Priority levels, queueing, fairness, rate limits, quotas.
  6. Deduplication
    • Idempotency keys, content-based dedup, time windows.
  7. Retries and Failure Handling
    • Backoff, dead-letter queues, poison-pill handling, fallback channels, circuit breaking.
  8. Scaling Strategies
    • Partitioning, horizontal scaling, multi-region, autoscaling triggers.
  9. Monitoring and Alerting
    • SLIs/SLOs, metrics, logs, traces, runbooks.
  10. Cost Considerations
  • Unit economics by channel, routing, batching, budgets, frequency caps.

State reasonable assumptions where needed and explain trade-offs.

Clarifying Questions to Ask Guidance

  • Clarify users, core use cases, read/write patterns, scale, latency, availability, and data retention.
  • State explicit assumptions before making sizing or architecture decisions.
  • Prioritize the functional path first, then address reliability, security, observability, and rollout.

What a Strong Answer Covers Guidance

  • A scoped requirements summary with concrete non-goals and success metrics.
  • API, data model, architecture, consistency, capacity, and operations.
  • Reasoned trade-offs among simple and scalable designs, including bottlenecks and failure modes.
  • A validation, monitoring, migration, and launch plan appropriate for the risk level.

Follow-up Questions Guidance

  • What breaks first at 10x traffic or data volume?
  • How would you degrade gracefully during dependency failures?
  • What metrics and alerts would prove the design is healthy after launch?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...