Kafka vs RabbitMQ: Architectural Differences for System Design Interviews

Kafka vs RabbitMQ for system design interviews: log vs broker architecture, routing, ordering, replay, and when RabbitMQ vs Kafka wins for each design.

Author: PracHub

Published: 4/13/2026

Kafka vs RabbitMQ: Architectural Differences for System Design Interviews

April 13, 2026
13 min read

Quick Overview

Kafka is a partitioned append-only log with consumer-owned offsets; RabbitMQ is a routing broker that deletes messages on ack. This guide works through the architectural differences interviewers actually probe — routing vs throughput, ordering guarantees, delivery semantics, and replay — grounded in real questions asked at Meta, DoorDash, Coinbase, Optiver, and Roblox.

Software EngineerFree

Kafka and RabbitMQ answer the same interview sentence — "these services talk asynchronously" — with opposite architectures. Kafka is a replicated append-only log: the broker stores events for a retention window, consumers track their own read position, and nothing is deleted because somebody read it. RabbitMQ is a routing broker: an exchange decides which queues each message lands in, a consumer acknowledges it, and the broker deletes it. Nearly every follow-up you will get in a system design interview (ordering, replay, fan-out, throughput) falls out of that one difference.

Key Takeaways

  • Kafka consumers pull from an immutable log and own their offset, so replay is native. RabbitMQ deletes on ack, so replay does not exist in a stock setup.
  • RabbitMQ routes in the broker with exchanges and bindings. Kafka gives you partitions; any routing beyond "same key, same partition" happens in your consumers.
  • Both are at-least-once by default. Neither removes the need for idempotent consumers, and saying so unprompted is a senior signal.
  • Kafka orders per partition, not per topic. RabbitMQ orders per queue, and only until competing consumers or redeliveries break it.
  • Rule of thumb: task queues with routing, priorities, or message TTLs point to RabbitMQ. Event streams several systems read independently, or anything needing replay, point to Kafka.

Video companion: a second pass on the log-vs-broker distinction and where event streaming diverges from event sourcing.

Kafka is a log, RabbitMQ is a broker: the difference everything else falls out of

Asked at Coinbase — Design real-time stock price viewer Build a system that shows millions of concurrent users the current price and best bid/ask for the symbols on their watchlists, fed by multiple upstream market data feeds. There are no historical queries, but the fan-out is enormous and several internal systems need to consume the same feed at the same time.

This problem has the log shape: one firehose of events, many independent readers. In Kafka, producers append records to a topic, and each topic is split into partitions. A partition is an ordered, immutable sequence of records on disk; every record gets an integer offset. The broker does not track who has read what. Each consumer group stores its own offset per partition, and Kafka retains records for a configured time or size window regardless of whether anyone read them.

That single design choice is why the Coinbase problem is comfortable on Kafka. The WebSocket gateways, the monitoring pipeline, and an aggregation service all read the same partitions at their own pace without interfering with each other. Adding a new consumer next quarter costs nothing at publish time: it joins as a new group, starts at offset 0 (or the tail), and backfills from retained data.

kafka vs rabbitmq architectural differences for system design interviews

RabbitMQ inverts every one of those properties. The broker owns delivery: it pushes a message to a consumer, marks it unacked, and deletes it once the consumer acks. Fan-out to N readers means binding N queues to one exchange, and only the queues already bound when a message arrives receive a copy. A reader added later receives nothing that was published before its queue existed.

The common mistake is treating these as two brands of the same product. They are different data structures. When an interviewer hears you say "Kafka is a log, RabbitMQ is a queue with a router in front," the follow-ups get easier, because you can derive the answers instead of memorizing them.

Routing: exchanges and bindings vs partitions and keys

Asked at DoorDash — Design a notification system A multi-tenant platform that delivers email, SMS, and push notifications, supporting both real-time and scheduled sends, per-user preferences, and quiet hours. The reliability requirements name deduplication, idempotency, and retries explicitly, and each channel needs its own handling.

Routing is where RabbitMQ earns its keep. A publisher sends a message to an exchange with a routing key like notify.sms.high, and binding rules decide which queues receive it. A topic exchange matches patterns (notify.sms.*, notify.*.high), a direct exchange matches exact keys, a fanout exchange copies to every bound queue. Dead-letter exchanges, queue TTLs, and priority queues come from broker declarations, and per-message expiration is a property you set when publishing; none of it is code you write in a consumer. For the DoorDash problem, channel-specific queues with a retry-then-dead-letter topology is a clean, defensible answer.

kafka vs rabbitmq architectural differences for system design interviews

Kafka's broker never inspects message content. The only routing primitive is the partition key: records with the same key land on the same partition. Anything richer ("SMS handlers should only see SMS events") becomes either a topic per channel or consumer-side filtering. At three channels that is fine. At forty routing rules you are rebuilding a topic exchange in application code, and a good interviewer will ask why you didn't use the broker that ships one.

The flip side is throughput. RabbitMQ pays for its routing with per-message bookkeeping: every message has delivery state, an unacked set, redelivery flags. Kafka does sequential appends and batched reads against an OS page cache, with no per-message broker state at all. That is the structural reason Kafka's throughput ceiling is far higher on the same hardware, not some tuning trick. If the interviewer's problem says "100k events per second, multiple downstream consumers," the log wins before you discuss configuration.

Ordering guarantees live in different places than candidates think

Asked at Meta — Design a real-time messenger A WhatsApp-style messenger supporting 1:1 and group chats, multi-device sync, receipts, and media, with explicit follow-ups on ordering and delivery guarantees at Meta scale. Every participant and every device must see a conversation's messages in the same order.

"Kafka guarantees ordering" is the most common half-truth in this comparison. Kafka guarantees order within a partition. Across partitions of the same topic there is no ordering at all. For the messenger, that means keying by conversation ID so all messages of one chat serialize through one partition, and accepting the consequence that a single conversation's throughput is capped at one partition's rate. With the idempotent producer enabled (the default in modern clients), that per-partition order survives producer retries too.

RabbitMQ's story sounds simpler and is more fragile. A queue is FIFO, so with exactly one consumer you get strict order. Add competing consumers and their processing interleaves. Have a consumer nack-and-requeue a message and it comes back while later messages, already prefetched by or in flight to other consumers, are being processed ahead of it. "RabbitMQ preserves order" is true precisely for one queue, one consumer, zero redeliveries — a configuration that also forfeits parallelism.

The senior move in the interview: name where the guarantee comes from (partition key vs queue discipline), name what breaks it (cross-partition reads vs competing consumers), and name the cost of keeping it (hot partitions vs single-threaded consumption).

Delivery semantics: at-least-once is the honest default for both

Asked at Optiver — Design a subscription push service An object-oriented pub/sub service with subscribe, unsubscribe, publish, and a client acknowledgement callback. The spec demands at-most-once delivery per published item and requires that an unsubscribe stops all future deliveries, including items already in flight.

Delivery semantics come from where you place the acknowledgement relative to the work, and that logic is identical in both systems. Commit the offset (Kafka) or ack (RabbitMQ) before processing and you get at-most-once: a crash after the ack loses the message. Process first and ack after, and you get at-least-once: a crash between the two redelivers, and your consumer sees duplicates. The delivery-semantics wording is the entire Optiver question, and its at-most-once demand is unusual; most production systems choose at-least-once and make the consumer idempotent, because losing a payment event is worse than seeing it twice.

Exactly-once deserves precision, because interviewers use it as a trap. Kafka's transactions give exactly-once within Kafka — a consume-transform-produce pipeline from one topic to another. The moment your consumer calls a payment gateway or an email API, you are back to at-least-once plus idempotency keys, no matter which broker you run. How to build that consumer correctly is its own topic: One Order, Two Charges: Idempotency in SQS and Kafka works the duplicate-charge scenario end to end, and "Just Make It Idempotent" Is Half an Answer covers what interviewers expect beyond the phrase itself.

When each wins: replay and reader count decide it

Asked at Meta — Design a price tracking system Ingest product URLs, crawl prices politely on a schedule across tens of millions of offers, store the price history, and alert users on drops or restocks. Each crawl is a discrete fetch job that should execute once, with retries and per-domain rate limits.

That crawler is a work queue even inside a data-heavy system: jobs, not events. Each unit of work is consumed once, retried on failure, dead-lettered when poisoned, and nothing downstream ever needs to re-read a completed fetch job. The results of the crawls are a different shape: price-change events feeding alerting and history, which several systems consume independently. Many strong designs use both, and naming that boundary is worth more than picking a side.

Two questions settle most broker choices faster than any feature checklist:

How many independent readers, now and later? One pool of workers splitting jobs is a queue. Several systems each needing every event is a log.

Will anyone need to re-read the past? Audit, backfill, reprocessing after a bug, bootstrapping a new service from history: any yes means Kafka. Once RabbitMQ acks a message it is gone; there is nothing to replay. The billing-grade recount in the Ad Click Aggregator design is a worked example of a requirement that is only satisfiable because the events still exist.

Two adjacent decisions are worth naming out loud. Before picking a broker at all, be sure the boundary belongs there; Synchronous vs Asynchronous: Where the Boundary Goes covers that earlier call. And RabbitMQ pushes to consumers (throttled by a prefetch limit) while Kafka consumers pull in batches, which puts backpressure control in the consumer's hands — Push vs Pull Architecture treats that trade-off in full.

RabbitMQ vs Kafka at a glance

DimensionRabbitMQKafka
Core abstractionQueues fed by exchangesPartitioned append-only log
Who tracks consumptionBroker, via per-message acksConsumer, via offset per partition
Message lifetimeDeleted after ackRetention window, independent of reads
ReplayNone once ackedNative — reset the offset
RoutingIn the broker: direct, topic, fanout, headersKey → partition; anything richer is consumer code
OrderingPer queue, single consumer, no redeliveriesPer partition, with idempotent producer
Fan-out to N readersN queues, each bound before messages flowN consumer groups, added any time
Consumer modelPush with prefetch limitPull with batched fetches
Delivery defaultAt-least-onceAt-least-once; exactly-once only Kafka-to-Kafka
Sweet spotTask queues, priority/TTL/dead-letter workEvent streams, analytics, audit and replay

The probe to expect: "Which one for order processing, and why?"

Asked at Roblox — Design a scheduled payment system Users schedule a payment for a future time, query its status, and cancel it before execution. Payments run through an external gateway, and the system must behave correctly across regions at scale — which makes execution-time delivery semantics and duplicate suppression the heart of the design.

This probe is a trap for either/or answers, because order processing contains both shapes.

Executing due payments is a work queue: each job runs once, gets retried with backoff on gateway timeouts, and lands in a dead-letter queue when it keeps failing. That is RabbitMQ's home turf (SQS competes for the same slot). The record of what happened is an event stream: the ledger, notification service, analytics, and reconciliation all consume payment events, and a billing dispute six weeks later means replaying history — Kafka's home turf. Key the event stream by payment ID and per-payment ordering comes with it.

Two additions separate strong answers here. First, the idempotency key at the gateway boundary is mandatory regardless of broker, because execution is at-least-once the moment retries exist. Second, cancellation is not a broker feature: a cancel racing an execution is settled by a state check in the database at execution time, not by trying to yank a message out of a queue. Pointing at the database instead of the broker for that problem shows you know what brokers are for.

Practice these on PracHub

Each of these turns one section of this page into a live interview defense. More in the full question bank.

FAQ

Is Kafka faster than RabbitMQ?

For sustained throughput, yes, and by design: sequential log appends and batched pulls beat per-message routing and ack bookkeeping. For single-message latency at modest volume, RabbitMQ is often lower: it pushes the instant a message arrives, while Kafka clients typically trade a little latency for batching (linger and minimum-fetch settings). In an interview, say which axis the problem cares about before quoting either claim.

Can RabbitMQ replay messages like Kafka?

Classic RabbitMQ queues cannot: an acked message is deleted. RabbitMQ Streams (added in 3.9) bolt a log abstraction with offsets onto the broker, but if replay is central to the design, interviewers expect the log-native tool rather than a queue with a log strapped on.

Which should I choose in a system design interview?

Map the requirement words. "Route by type/priority," "process each job once," "delay and retry" point to RabbitMQ. "Multiple consumers of the same events," "replay," "audit," "stream processing," "very high throughput" point to Kafka. Saying the mapping out loud is the answer; the brand name is secondary.

Does Kafka give exactly-once delivery?

Within Kafka, yes: the idempotent producer plus transactions make a consume-transform-produce pipeline exactly-once from topic to topic. End to end with an external system — a payment gateway, an email provider — no broker gives exactly-once; you get at-least-once plus an idempotent consumer.

Can I use Kafka and RabbitMQ in the same system?

Yes, and many production systems do: RabbitMQ (or SQS) for work distribution with retries and dead-lettering, Kafka as the durable event backbone that fan-out consumers and replay depend on. In an interview, the value is naming where one hands off to the other and why.

Is Kafka push or pull?

Kafka consumers pull batches from the broker, which puts backpressure control in the consumer: it fetches only when ready. RabbitMQ pushes to consumers and throttles via the prefetch (QoS) limit. Neither is strictly better; pull favors throughput and consumer pacing, push favors low idle latency.


Comments (0)