Async Services, Idempotency, And Retries
Asked of: Software Engineer
Last updated
What's being tested
Interviewers are probing your ability to design robust asynchronous services that tolerate network failures using correct idempotency and retry strategies while preserving correctness, auditability, and performance. Expect to justify tradeoffs between at-least-once vs exactly-once behaviors, pick storage and deduplication patterns, and describe operational controls (timeouts, backoff, metrics). Databricks values systems that remain correct under partial failure and that make clear, testable guarantees a software engineer can implement and maintain.
Core knowledge
-
Idempotency — guarantee that repeating the same operation has the same effect; implement via an idempotency key (client-generated or server-generated) stored with outcome and TTL to deduplicate retries atomically.
-
At-least-once vs exactly-once — at-least-once is simpler (retries may duplicate) while exactly-once requires coordination (dedup stores, distributed locks, or transactional sinks) and costs latency and complexity.
-
Deduplication store patterns — use a single-row unique constraint in
Postgresor a write-once record inRedis/Cassandrato record(idempotency_key, status, result); keep TTL to bound storage growth and support eventual GC. -
Retry policies & backoff — prefer exponential backoff with jitter to avoid thundering herds; cap retries and expose a failure mode to callers after N attempts. Quantify: start 100–200ms, double to a max ~5–10s.
-
Synchronous vs asynchronous flows — synchronous (blocking) authorization provides immediate result but ties up clients; asynchronous (webhook, polling, callback) scales better for slow or multi-hop operations like external authorization.
-
Compensating actions & reversals — for non-idempotent side effects (payments, stock reservation), model operations as two-phase: tentatively reserve then commit/settle, or create a compensating transaction (reversal) when downstream fails.
-
Message delivery semantics — with
Kafka/SQSexpect at-least-once delivery; design consumers to be idempotent or use an exactly-once stream processor (e.g.,KafkaStreams with EOS) when necessary. -
Ordering and causality — if order matters, use partitioning keys and monotonic offsets; for multi-participant operations, use logical timestamps or vector clocks to reason about concurrent retries and deletions.
-
Durability and audit — persist the original request, idempotency key, response, and timeline (timestamps) to support reconciliation, dispute resolution, and retries. Use append-only logs for audit trails.
-
Performance tradeoffs — deduplication lookups add latency (one extra read); choose read-before-write vs conditional write (
INSERT ... ON CONFLICT DO NOTHING) depending on load: conditional writes scale better under high contention. -
Failure modes to test — network timeouts, duplicate client retries, partial downstream success, and long-tail latency spikes (
p99behaviour). Define SLAs: e.g., respond within 500ms for sync paths, otherwise escalate to async. -
Security and id generation — prefer collision-resistant ids (UUIDv4 or client nonce with HMAC) as idempotency keys; authenticate and bind keys to client credentials to prevent replay across accounts.
Worked example — Design a Visa-Style Card Payment Processing System
First 30 seconds: clarify scope (merchant-side acquirer vs global network), required latency for authorization, failure semantics (is double-charge acceptable?), and which parties can retry (merchant, gateway, issuer). State assumptions: we design an acquirer gateway handling authorizations and routing to issuers, with a requirement to avoid duplicate captures.
Skeleton pillars: (1) API & idempotency: require client-supplied idempotency_key per payment attempt and store (key, status, response) atomically; (2) Routing & async authorization: send authorization to issuer synchronously with a bounded timeout, else mark as PENDING and enqueue for async retry; (3) Settlement & ledger: an append-only settlement ledger records captures and reversals, with unique constraints to prevent duplicate entries; (4) Reconciliation & audit: background jobs reconcile pending states with issuers and emit compensating transactions for timeouts.
One tradeoff to flag: synchronous blocking gives merchant immediate truth but creates load and failures on issuer slow paths; asynchronous decouples latency but complicates customer UX and requires clear guarantees (e.g., eventual confirmation within X minutes). Implementation detail: use INSERT ... ON CONFLICT to create idempotency records atomically and return stored response if conflict occurs.
Close by stating next steps: if more time, add fraud filtering pipeline, a test harness to simulate partial issuer failures, and dashboards for p99 latency, retry counts, and reconciliation gaps.
A second angle — Design Chat APIs, Storage, and Message Flows
Chat emphasizes ordered, low-latency delivery and offline recovery, so idempotency focuses on message deduplication and reconnection correctness. Use client-generated monotonic message_id or UUID plus a send-sequence per conversation; server stores (conversation_id, message_id) to reject duplicates. For deletions and edits, model operations as idempotent commands with tombstones and causal metadata so offline clients converge (apply ops in sequence or use CRDTs for idempotent merge). Retries: clients should retry sends to server until acknowledged, server must be able to reply with the canonical message id/result to any duplicate send. The core concept—protecting side effects from duplicate retries—remains identical, but constraints (ordering, per-conversation partitioning, reconnect reconciliation) change data model and storage choices.
Common pitfalls
Pitfall: Treating retries purely as a client problem.
Many engineers assume client retries alone solve transient failures. In reality the server must detect duplicates (idempotency store) and return the prior result; otherwise clients will see inconsistent outcomes and operators face reconciliation headaches.
Pitfall: Over-engineering exactly-once when at-least-once suffices.
Chasing full exactly-once semantics across multiple systems often costs latency and complexity; prefer idempotent consumer/producer patterns and compensating transactions unless business correctness absolutely requires atomic cross-system commit.
Pitfall: Forgetting TTLs and GC for idempotency records.
Keeping deduplication records forever prevents replays but creates unbounded storage. Define an idempotency window aligned with business semantics (e.g., 24–72 hours) and a safe GC policy to reclaim space.
Connections
Interviewers may pivot to distributed transactions / two-phase commit, event sourcing / CQRS, or observability (tracing retries end-to-end). Be ready to discuss how idempotency integrates with monitoring (retry_count, duplicate_rate, reconciliation lag).
Further reading
-
Designing Data-Intensive Applications — Martin Kleppmann — strong background on consistency, ordering, and deduplication patterns.
-
Stripe: Idempotency Keys — practical implementation patterns and tradeoffs for payment systems.
Practice questions
- Lowest-Price Book Aggregator Fanning Out to Hundreds of Async BookstoresDatabricks · Software Engineer · Onsite · medium
- Design a Visa-Style Card Payment Processing SystemDatabricks · Software Engineer · Onsite · medium
- Design Chat APIs, Storage, and Message FlowsDatabricks · Software Engineer · Onsite · medium
- Design Messaging With Message DeletionDatabricks · Software Engineer · Onsite · medium
Related concepts
- Idempotent API DesignSystem Design
- API Idempotency And Concurrency ControlSystem Design
- Idempotency And Concurrency ControlSystem Design
- Fault Tolerance, Idempotency, And Concurrency ControlSystem Design
- Distributed Systems Correctness And IdempotencySystem Design
- Idempotency, Deduplication, and Delivery SemanticsSystem Design