Webhook vs API: Who Calls Whom, and What Interviewers Actually Probe
Quick Overview
An API is a request you make; a webhook is a callback the other system makes to you, and that inversion drives everything else: polling cost at real request volumes, at-least-once delivery with retries and dead-letter queues, HMAC signature verification with replay windows, and out-of-order events. This guide teaches each layer through real company-tagged interview questions from DoorDash, Microsoft, Atlassian, Soti, Attentive, and Mistral AI, plus the delivery and signature schemes that Stripe-style integration rounds grade on.
You call an API; a webhook calls you. That is the whole distinction. An API is a request/response interface your code invokes when it wants something. A webhook is an HTTP callback the other system fires at an endpoint you expose, when something happens on their side. Everything that follows (polling costs, retry semantics, signature verification, out-of-order delivery) falls out of that one inversion of who initiates the connection. Interviewers know this, which is why "webhook vs API" questions are rarely definitional. They are systems questions wearing a vocabulary costume: the interviewer asks "how would the client find out the payment settled?" and watches whether you reach for polling, a callback, or a socket, and whether you can defend the choice with numbers.
Key Takeaways
- A webhook is not an alternative to an API. It is an API call, in the reverse direction. The real comparison is push (webhooks) vs pull (polling), and the deciding variables are event rate, acceptable staleness, and who pays for wasted requests.
- Polling traffic scales with client count × poll frequency, not with how often anything changes. 10,000 clients polling every 10 seconds is 86.4M requests a day even on a day when nothing changed.
- A bare webhook is at-most-once. Production delivery means a durable queue, exponential backoff with jitter, and a dead-letter queue. That upgrade buys you at-least-once, which forces idempotent consumers.
- Secure a webhook endpoint with an HMAC signature computed over the timestamp plus the raw body, verified with a constant-time compare, inside a short replay window. "It's HTTPS" is not an answer.
- Never trust webhook ordering. Carry a version or sequence number, or send a thin payload and have the consumer fetch current state.
The difference is who initiates — and who therefore carries the burden
Asked at DoorDash — Design a resilient bootstrap API On launch, a client app makes one call that has to return all the data its first screen needs. Behind that single endpoint the data is owned by three different internal services, so your job is an aggregator: fan out to all three, compose one unified response, and stay sensible when a dependency is slow or down.
This is the API model in its purest form. The client knows exactly when it needs data (app launch), so it asks, once, and the server answers. The server keeps no memory of the client's interest. No subscription table, no delivery worker, no retry queue. When the response is sent, the transaction is over. That statelessness is why request/response is the default for almost everything: it is the cheapest contract two systems can have.
The model breaks on one specific question: how does the client find out something changed? Request/response gives you two bad options. Ask constantly, and you burn traffic asking about things that did not happen. Ask rarely, and you accept staleness. Webhooks resolve the dilemma by inverting the arrow: the client registers a URL once, and the server calls it when, and only when, there is news.
That inversion is not free, and this is the point candidates miss. The subscriber must now run a public, always-available HTTP endpoint, and the provider must now run delivery infrastructure: subscription storage, workers, retries, monitoring. A webhook is not a "faster API." Per request it is the same HTTP POST. What it eliminates is the requests that would have returned nothing.
If you want the general framework for reasoning about this trade-off beyond webhooks (feeds, caches, notification fan-out), the push vs pull architecture guide covers it; this article stays on the webhook-specific mechanics.
Polling turns "no news" into your biggest traffic source
Asked at Soti — Design an IP blacklist API Build the API and backend for a firewall blacklist. Security admins and automated abuse detectors add and remove banned IPs and CIDR ranges, while fleets of gateways query the service at very high throughput to decide whether each incoming request should be blocked. The interviewer pushes hardest on cache design — and on how those gateways learn about changes.
The gateways obviously cannot call the API once per packet, so they cache the blacklist locally. Now you have an invalidation problem, and the naive answer is a poll loop: every gateway re-fetches (or asks "changed since T?") on an interval.
Do the arithmetic out loud in the interview, because it is the whole argument:
- 10,000 gateways polling every 10 seconds = 1,000 requests/second = 86.4M requests/day.
- Suppose the blacklist actually changes 500 times a day, spread out — one change roughly every three minutes, so each lands in a different poll window. Then every gateway sees about 500 non-empty polls a day: ~5M polls fleet-wide return a delta, and the other ~81.4M (about 94%) return "nothing changed." The overwhelming majority of your traffic exists to discover the absence of events.
- And you still eat staleness: with a 10-second interval, an attacker blacklisted at t=0 gets up to 10 more seconds of open door, ~5 seconds on average.
Push keeps the useful 5M and deletes the empty 81.4M. The blacklist service keeps a registry of gateway callback endpoints (or a pub/sub channel) and delivers a delta on each change: 500 changes × 10,000 gateways is the same 5M data-carrying messages, now with nothing wasted around them — 94% less traffic, and propagation latency drops from seconds to milliseconds. Traffic now scales with the event rate, which is the quantity you actually care about.

Say when polling wins, because it sometimes does, and naming the boundary is the senior signal:
- Few clients, or events on nearly every poll. If the answer changes every interval anyway, polling wastes nothing and is drastically simpler.
- The client cannot expose an endpoint. Browsers, mobile apps, and services behind NAT or a corporate firewall cannot receive an inbound POST.
- You need reconciliation anyway. Webhooks get lost; a periodic full sync is the safety net, and if you must build the poll path regardless, at low scale it may be the only path you need.
That second constraint is common enough to deserve its own map, because a browser tab still wants events; it just cannot run a public server. What actually differs across the channels:
| Channel | Who initiates | Notification latency | Consumer needs a public endpoint | Wasted requests | Fits best |
|---|---|---|---|---|---|
| Polling | Consumer | avg. half the poll interval | No | High when events are rare | Simple integrations, reconciliation, low client counts |
| Webhook | Producer | ~milliseconds | Yes | None | Server-to-server event delivery |
| Long-polling | Consumer | ~milliseconds | No | Low (held connections instead) | Near-real-time without inbound access |
| SSE | Consumer opens, server streams | ~milliseconds | No | None after connect | Server→browser feeds: dashboards, notifications |
| WebSocket | Consumer opens, both send | ~milliseconds | No | None after connect | Bidirectional: chat, collaborative editing, games |
Long-polling is polling with the waste squeezed out: the server holds the request open until an event arrives or a timeout hits, so latency approaches webhook levels while the consumer stays behind its firewall. The trade is a held connection per waiting client. SSE keeps one HTTP connection open for a one-way, auto-reconnecting event stream into a browser. WebSockets upgrade to a full-duplex channel and are the answer when the client talks back at high frequency.
The interview heuristic: webhooks between servers, sockets or SSE to end-user devices, polling as the reconciliation floor under either. Real systems layer them. A payment provider webhooks your backend, your backend pushes over SSE to the user's open tab, and a nightly sync re-polls to catch anything both layers dropped.
Async jobs: return 202, then decide who waits
Asked at Mistral AI — Design a PDF-to-Markdown Inference API Compose two existing building blocks — a CPU-heavy function that splits and rasterizes PDF pages, and a GPU-heavy OCR engine — into a production inference service that turns uploaded PDFs into Markdown. Jobs take seconds to minutes, so the interface has to handle work that outlives any single HTTP request.
This is where webhook-vs-API stops being either/or and becomes a design decision inside one API. A 40-page PDF might take 90 seconds of GPU time; you cannot hold the HTTP connection open for that. So the submit endpoint returns 202 Accepted with a job ID, and now you must choose how the caller learns the job finished:
- Client polls
GET /jobs/{id}. Simplest possible contract. The cost is latency (with a 5-second poll interval, the client learns of completion 2.5 seconds late on average) plus the wasted-request math from the previous section, multiplied across every in-flight job. - Client registers a
callback_urlat submit time. Zero wasted requests, near-zero notification latency. The cost lands on both sides: the client must run a reachable endpoint, and you must build the delivery pipeline in the next section. Most of the provider-side machinery in this article exists to support this option. - Both. Webhook as the latency optimization, polling as the fallback and reconciliation path.
Option 3 is the answer that reads as production experience, if you justify it: webhook deliveries fail — the subscriber deploys at the wrong moment, their cert expires, a proxy eats the request — so the poll endpoint is not redundant, it is the source of truth the callback merely accelerates. This "status endpoint is authoritative, notification is a hint" framing will carry you through most integration-round follow-ups. For structuring the rest of the answer (resource naming, status codes, versioning), use the API design interview framework.
A bare webhook is at-most-once; production delivery is a pipeline
Asked at DoorDash — Design an alert notification system Design multi-tenant alerting for operational incidents: monitoring sources emit events, users configure routing rules and escalation chains, and the platform must reach responders over email, SMS, push, chat integrations, and phone — at millions of events a day, with p95 time-to-first-notification under 30 seconds and at-least-once delivery.
The naive webhook implementation is an HTTP POST inside the request handler that produced the event. Enumerate how it dies, because interviewers score you on named failure modes:
- Subscriber down for 90 seconds during a deploy. Every event in that window is gone forever.
- Endpoint returns 500. Same result, unless you retry, which the naive version does not.
- Endpoint is slow. A subscriber taking 10 seconds per request now stalls whatever thread sent the event. Their latency became your latency.
- One dead subscriber, many healthy ones. Without isolation, retries for the dead one consume the capacity meant for everyone else. Head-of-line blocking, cross-tenant.
The fix is the same shape at Stripe or any serious webhook provider:

The load-bearing decisions, with numbers:
- Persist before acknowledging. The event is written to durable storage before the producer gets its 200. Everything downstream can crash and recover; nothing is lost between "it happened" and "it was queued."
- Treat timeout and non-2xx identically: failed. Set the timeout at 5–10 seconds. A subscriber that streams a response for 60 seconds is a failure, not a slow success.
- Exponential backoff with jitter. A schedule like 1m, 5m, 30m, 2h, 12h avoids hammering a subscriber that is down because it is overloaded, and jitter prevents every failed delivery from retrying in the same second. Stripe retries failed webhook deliveries with increasing delays for up to about three days before giving up; that number is a useful anchor when an interviewer asks "retry for how long?"
- Dead-letter, don't drop. Exhausted deliveries go to a DLQ with the full payload and attempt history, alert an operator, and support replay. Auto-disable subscriptions that have failed continuously for days, and notify the owner. Silently retrying into a long-dead endpoint forever is its own bug.
- Per-subscriber isolation. Partition the queue (or maintain per-subscriber sub-queues) so one broken endpoint cannot starve the rest. Choosing the broker that gives you those partitions is its own interview topic; the Kafka vs RabbitMQ comparison covers which primitives each one actually provides.
Name the counterexample too, because not every provider builds this pipeline: GitHub webhooks do not automatically retry a failed delivery. The failure is recorded, and you redeliver manually or through the redelivery API. That is exactly why serious GitHub integrations run their own reconciliation poll, and why "what does this provider actually do on failure?" is the first question to ask about any webhook you consume.
Notice what happened to the delivery guarantee. Fire-and-forget was at-most-once. Persist-and-retry is at-least-once — the retry that fixes lost deliveries also creates duplicate deliveries, whenever a subscriber processes the POST but the 200 never makes it back. Exactly-once delivery over plain HTTP is not on the menu. That hands the next problem to the consumer.
At-least-once delivery makes idempotency the consumer's problem
Asked at Atlassian — Design a tagging system and REST APIs Design a tagging system with a strong REST focus: create, rename, and delete tags; attach and detach them from entities like pages and issues; search entities by tag with AND/OR semantics. The non-functional requirements call out concurrency and idempotency on attach/detach explicitly — the interviewer wants to see you handle the same request arriving twice.
Every event a provider sends must carry a stable, unique ID (evt_9f3k…), and every consumer must use it to make redelivery harmless. The cleanest implementation rides on a unique constraint:
# Pseudocode — adapt transaction/rowcount details to your driver or ORM.
def handle_event(event):
with db.transaction():
inserted = db.execute(
"INSERT INTO processed_events (event_id) VALUES (%s)"
" ON CONFLICT DO NOTHING",
(event["id"],),
).rowcount
if inserted == 0:
return # duplicate delivery — already handled, ack and move on
apply_side_effects(event) # same transaction as the marker
Two details separate a working version from a résumé version:
- The marker and the side effects must commit atomically. Record the ID first and apply effects after the transaction, and a crash in between means the retry gets deduplicated against work that never happened — you traded duplicate processing for silent loss. Same transaction, or an outbox pattern.
- The dedup store must outlive the provider's retry horizon. A Redis set with a 1-hour TTL loses to a provider that retries for three days; delivery #6 on day two sails through as "new." Size the retention window to the sender's documented retry schedule, plus margin for manual replays from their DLQ.
Idempotency is a deeper topic than one dedup table: natural-key writes, idempotency keys on the API side, the difference between deduplicating a message and making an operation idempotent. One Order, Two Charges works the queue-consumer version in detail, and "Just make it idempotent" is half an answer covers the follow-up questions interviewers ask after you say the word.
Anyone on the internet can POST to your endpoint
Asked at Microsoft — Design a Secure Copilot API Design the API for a multi-tenant enterprise AI copilot: authenticated users send prompts, retrieve responses, and invoke tenant-specific tools against tenant-private data. The round is explicitly a security interview — authentication, authorization, token lifecycle, and abuse prevention are the core of the exercise, not an afterthought.
API security has a mature playbook: OAuth flows, scoped tokens, tenant isolation, rate limits (the rate limiting algorithms guide covers that layer). Webhook receivers need a different playbook, because the roles flipped: your endpoint is a public URL, and anyone, not just the provider, can POST convincing-looking JSON at it. An unauthenticated webhook receiver is an unauthenticated write API into your system.
The standard answer is a shared-secret HMAC, verified the way Stripe does it:
import hmac, hashlib, time
TOLERANCE = 300 # seconds — the replay window
def verify(raw_body: bytes, sig_header: str, ts_header: str, secret: bytes) -> bool:
if abs(time.time() - int(ts_header)) > TOLERANCE:
return False # stale: replayed or badly skewed clock
signed_payload = ts_header.encode() + b"." + raw_body
expected = hmac.new(secret, signed_payload, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, sig_header)
Each line exists because of a specific attack, and interviewers probe for whether you know which:
- Sign the timestamp together with the body. Signing the body alone lets an attacker capture one valid delivery and replay it forever. Binding the timestamp into the signature plus a 5-minute tolerance window shrinks replays to a window you can additionally dedup with the event ID. GitHub's
X-Hub-Signature-256is the instructive counterexample: it is HMAC-SHA256 over the body only, no timestamp, so replay defense falls entirely on your own event-ID dedup. Naming that difference is a strong answer to "what does the timestamp buy you?" compare_digest, never==. String equality short-circuits at the first differing byte, and the timing difference leaks how much of a forged signature is correct.- Verify the raw bytes. If your framework parses JSON and you re-serialize to check the signature, key ordering or whitespace changes break verification intermittently — a classic "works locally, fails for one customer" bug. Read the body before any middleware touches it.
- Per-subscriber secrets, rotated with overlap. During rotation the sender signs with the new secret while receivers accept either, then the old one is retired. One global secret means one leak compromises every subscriber.
Stripe's integration round tests exactly this material against their real API — signature verification, replay handling, idempotent retries — and the Stripe integration round guide walks through the edge cases they grade on.
Webhooks arrive out of order; design like it
Asked at Attentive — Design Broadcast Scheduling APIs Design a broadcast messaging service: each company has subscribers, a client schedules a message for a future time, scheduled broadcasts are queryable by time range, and the system must eventually deliver the message to every intended subscriber. "Eventually, to every subscriber" is doing a lot of quiet work in that sentence.
Once delivery involves retries and a parallel worker pool, ordering is gone. subscriber.updated fails and lands in the 5-minute retry bucket; subscriber.deleted, fired two seconds later, succeeds immediately. Your consumer sees delete-then-update and resurrects a record that should be gone. No provider that retries can promise global order, and most explicitly disclaim it.
Three defenses, in the order you should reach for them:
- Version or sequence numbers on the object. Each event carries the object's version; the consumer applies an event only if its version is newer than what is stored, and drops stale ones. Cheap, local, and it turns reordering from corruption into a no-op.
- Thin payloads. The event says what changed (
subscriber sub_123 changed), not the new state; the consumer calls back toGET /subscribers/sub_123for current truth. Out-of-order hints collapse harmlessly, and a burst of 50 events for one object becomes 50 cheap fetch-latest calls (or one, after debouncing). The costs: up to one extra read per event, less with debouncing, and the GET can race a subsequent write — acceptable, since you always end at current state, which is what you wanted. - Partition-ordered transport. If you control both sides, a log like Kafka partitioned by object key gives per-key FIFO. This is an in-house option; you cannot impose it on third-party webhook consumers over HTTP.
The thin-payload pattern also quietly fixes a security problem: if the payload carries no state, a forged or replayed delivery can only trigger a read.
Practice these on PracHub
Each of these is a real, company-tagged question from the PracHub question bank that drills one layer of this article:
- Design an alert notification system (DoorDash) — the full delivery pipeline: at-least-once, retries, escalation, and a 30-second p95 to hold while doing it.
- Design Broadcast Scheduling APIs (Attentive) — scheduled fan-out, "eventually every subscriber," and the ordering problems that follow.
- Design a scalable notification system (Airbnb) — event-triggered vs bulk delivery across email, SMS, and push; the multi-channel version of the pipeline.
- Design a PDF-to-Markdown Inference API (Mistral AI) — the 202-then-notify decision on a job that outlives the request.
- Design an IP blacklist API (Soti) — push vs poll for cache invalidation, with the throughput math front and center.
- Design a tagging system and REST APIs (Atlassian) — idempotent writes under concurrency, the consumer-side skill at-least-once delivery demands.
For more worked API rounds, the top 10 API design questions for 2026 and the broader resources library cover the adjacent topics interviewers chain into.
Comments (0)