Design a group chat system

Read the full interview experience this question came from →

Quick Overview

Airbnb's onsite system design round asks you to design the backend for a real-time group chat and messaging system: 1:1 and group conversations, near-real-time delivery, accurate unread counts, and full paginated history at roughly 1B messages a day and 5M concurrent connections. A complete answer covers the guarantees, APIs, data model and partitioning, write path and ordering, WebSocket fan-out and presence, offline delivery and multi-device sync, unread counts, membership changes, failure modes and observability. This walkthrough gives a corrected end-to-end reference answer for every part, including the trade-offs interviewers probe.

Design a group chat system

Company: Airbnb

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

##### Question Design the backend for a real-time **group chat / messaging system** at Airbnb scale — the in-app messaging that connects guests and hosts, including multi-party group conversations. Users create a conversation with one or more participants, send text messages, and see messages from others appear in near real time. Each participant sees an accurate unread-message count and can scroll back through the full history of any conversation. **Core features** - 1:1 and group conversations - Send and receive text messages (small payloads; attachments by reference) - Group membership management (create a group, add and remove participants) - Per-conversation unread counts and paginated history - Message history sync across a user's multiple devices - Offline delivery for participants who are not connected **Non-functional requirements** - High availability and low latency on both send and receive (sub-second p95 delivery to online participants) - Durability: a message the server has acknowledged is never silently lost - A clearly *defined* consistency and ordering guarantee — say what you require and why **Constraints and assumptions** (state your own numbers; these are reasonable targets) - ~50M monthly active users, ~5M concurrent connected users at peak - ~1B messages/day, roughly 12K writes/sec average and ~60K writes/sec at 5x peak - Most conversations are small-to-medium (up to ~100 members). Discuss explicitly what changes when a conversation grows to thousands or tens of thousands of members - Read-heavy: opening a conversation loads the latest ~30 messages; history scroll is paginated **Walk through an end-to-end design covering:** 1. **Requirements and guarantees** — the clarifying questions you would ask, and the ordering / delivery / consistency guarantees you commit to (per-conversation total order, at-least-once delivery with de-duplication, and so on). 2. **APIs** — conversation and membership endpoints, send-message, history fetch, and the delivery/read acknowledgement calls. 3. **Data model and storage** — the entities, their partition and clustering keys, and which storage engine you pick for the message log versus conversation metadata. 4. **Write path** — what happens when a user hits send: validation, sequence assignment, where the durability boundary sits, and when the sender is acknowledged. 5. **Ordering** — how a message gets its position in the conversation, and why wall-clock timestamps alone are not enough. 6. **Real-time fan-out and presence** — the connection layer, how you locate a recipient's live connection, and how routing scales to millions of concurrent sockets. 7. **Fan-out strategy** — fan-out on write versus fan-out on read, and where the crossover sits as group size grows. 8. **Offline delivery, retries, and multi-device sync** — durable inboxes or log replay, push wake-ups, reconnect/resync, and how a user's second device catches up. 9. **Unread counts and history pagination** — an accurate per-user unread count per conversation, and efficient deep scroll-back. 10. **Group membership changes** — what a newly added member can see, and whether a removed member receives in-flight messages. 11. **Read receipts and typing indicators** — optional, but discuss the cost and the trade-offs in large groups. 12. **Scaling, failure modes, and observability** — sharding, hot conversations, what breaks under each failure, and the metrics and alerts you would run on. ### Clarifying questions to ask - What is the maximum group size — bounded small groups (<= 100) or broadcast-scale channels? This decides whether fan-out on write is acceptable. - Do we need delivery and read receipts and typing indicators, or only message send plus unread counts? - Is per-conversation total order enough, or is cross-conversation global ordering required (almost never)? - What is the retention policy (forever versus TTL), and are edits, deletions, or recall in scope? - Is end-to-end encryption required, or is transport plus at-rest encryption sufficient? - Which platforms must we support? Mobile with offline sync and push notifications changes the design substantially. - For a 100-member group where 80 are offline, do offline members get a push notification plus sync-on-reconnect, or is real-time delivery best-effort for online users only? ### Follow-up questions - How do you guarantee a message is never lost if the server crashes after acknowledging the client but before fan-out completes? - How does the design change for a conversation with 100,000 members? - A user is removed from a group while a message is in flight — should they receive it, and how do you enforce that without a per-message, per-user database read?

Overview: Airbnb's onsite system design round asks you to design the backend for a real-time group chat and messaging system: 1:1 and group conversations, near-real-time delivery, accurate unread counts, and full paginated history at roughly 1B messages a day and 5M concurrent connections. A complete answer covers the guarantees, APIs, data model and partitioning, write path and ordering, WebSocket fan-out and presence, offline delivery and multi-device sync, unread counts, membership changes, failure modes and observability. This walkthrough gives a corrected end-to-end reference answer for every part, including the trade-offs interviewers probe.

Read the full Airbnb Software Engineer interview experience this question came from

|Home/System Design/Airbnb
Airbnb logo
Airbnb
Jan 19, 2026
hardSoftware EngineerOnsiteSystem Design
23
0
Question

Design the backend for a real-time group chat / messaging system at Airbnb scale — the in-app messaging that connects guests and hosts, including multi-party group conversations. Users create a conversation with one or more participants, send text messages, and see messages from others appear in near real time. Each participant sees an accurate unread-message count and can scroll back through the full history of any conversation.

Core features

  • 1:1 and group conversations
  • Send and receive text messages (small payloads; attachments by reference)
  • Group membership management (create a group, add and remove participants)
  • Per-conversation unread counts and paginated history
  • Message history sync across a user's multiple devices
  • Offline delivery for participants who are not connected

Non-functional requirements

  • High availability and low latency on both send and receive (sub-second p95 delivery to online participants)
  • Durability: a message the server has acknowledged is never silently lost
  • A clearly defined consistency and ordering guarantee — say what you require and why

Constraints and assumptions (state your own numbers; these are reasonable targets)

  • ~50M monthly active users, ~5M concurrent connected users at peak
  • ~1B messages/day, roughly 12K writes/sec average and ~60K writes/sec at 5x peak
  • Most conversations are small-to-medium (up to ~100 members). Discuss explicitly what changes when a conversation grows to thousands or tens of thousands of members
  • Read-heavy: opening a conversation loads the latest ~30 messages; history scroll is paginated

Walk through an end-to-end design covering:

  1. Requirements and guarantees — the clarifying questions you would ask, and the ordering / delivery / consistency guarantees you commit to (per-conversation total order, at-least-once delivery with de-duplication, and so on).
  2. APIs — conversation and membership endpoints, send-message, history fetch, and the delivery/read acknowledgement calls.
  3. Data model and storage — the entities, their partition and clustering keys, and which storage engine you pick for the message log versus conversation metadata.
  4. Write path — what happens when a user hits send: validation, sequence assignment, where the durability boundary sits, and when the sender is acknowledged.
  5. Ordering — how a message gets its position in the conversation, and why wall-clock timestamps alone are not enough.
  6. Real-time fan-out and presence — the connection layer, how you locate a recipient's live connection, and how routing scales to millions of concurrent sockets.
  7. Fan-out strategy — fan-out on write versus fan-out on read, and where the crossover sits as group size grows.
  8. Offline delivery, retries, and multi-device sync — durable inboxes or log replay, push wake-ups, reconnect/resync, and how a user's second device catches up.
  9. Unread counts and history pagination — an accurate per-user unread count per conversation, and efficient deep scroll-back.
  10. Group membership changes — what a newly added member can see, and whether a removed member receives in-flight messages.
  11. Read receipts and typing indicators — optional, but discuss the cost and the trade-offs in large groups.
  12. Scaling, failure modes, and observability — sharding, hot conversations, what breaks under each failure, and the metrics and alerts you would run on.

Clarifying questions to ask Guidance

  • What is the maximum group size — bounded small groups (<= 100) or broadcast-scale channels? This decides whether fan-out on write is acceptable.
  • Do we need delivery and read receipts and typing indicators, or only message send plus unread counts?
  • Is per-conversation total order enough, or is cross-conversation global ordering required (almost never)?
  • What is the retention policy (forever versus TTL), and are edits, deletions, or recall in scope?
  • Is end-to-end encryption required, or is transport plus at-rest encryption sufficient?
  • Which platforms must we support? Mobile with offline sync and push notifications changes the design substantially.
  • For a 100-member group where 80 are offline, do offline members get a push notification plus sync-on-reconnect, or is real-time delivery best-effort for online users only?

Follow-up questions Guidance

  • How do you guarantee a message is never lost if the server crashes after acknowledging the client but before fan-out completes?
  • How does the design change for a conversation with 100,000 members?
  • A user is removed from a group while a message is in flight — should they receive it, and how do you enforce that without a per-message, per-user database read?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...