Interview conceptSystem Design

Resilient APIs And External Dependency Integration

Asked of: Software Engineer

Last updated

What's being tested

Interviewers probe your ability to integrate and expose external services reliably: designing APIs and clients that tolerate flakiness, control resource usage, and fail predictably. Expect questions that check knowledge of timeouts, retries, idempotency, backpressure and isolation (bulkheads), and observable failure modes (p99 latency, success rate). Snowflake values engineers who can keep user-facing paths available and debuggable while respecting external SLA variability.

Core knowledge

  • Retries: implement exponential backoff with jitter (e.g., base b, attempts n → delay ≈ b * 2^n ± jitter) and a capped max attempts to avoid cascading load.

  • Idempotency: require idempotency keys for non-idempotent server actions; client SDK should attach Idempotency-Key and deduplicate on the server to avoid side-effect duplication.

  • Circuit breaker: use circuit breaker (closed/half-open/open) to short-circuit calls after threshold failures; combine with health checks and cooldown windows to prevent retry storms.

  • Bulkhead isolation: partition pools by external dependency or operation type (threadpool/connection pool) so one slow API cannot exhaust all resources; size pools using capacity planning.

  • Timeouts vs deadlines: use per-call timeout shorter than user-facing deadline, and propagate a deadline header so downstreams can fail fast; avoid infinite blocking.

  • Rate limiting: enforce token bucket or leaky bucket per-tenant and per-external-host; implement global and per-host quotas for crawlers or outward API calls.

  • Caching & staleness: use read-through caches with TTL and stale-while-revalidate for soft degradation; consider cache invalidation costs — LRU works up to ~10M keys; beyond that use approximate structures or sharding.

  • Backpressure & queues: buffer with bounded queues; apply admission control (reject or degrade) when queue length > threshold; Little’s law applies: L=λWL = \lambda W to size pools.

  • Observability: emit structured logs, metrics (request_rate, error_rate, p50/p95/p99), and distributed traces with a correlation id; alert on success-rate SLOs, not just latency.

  • Async fallback & DLQ: convert sync calls to async via queues for brittle dependencies; failed items go to a dead-letter queue for investigation and replay.

  • Security & token handling: cache third-party tokens, refresh proactively before expiry; on flaky token endpoints, use retry+Jitter and local validation (e.g., verify signatures) to reduce calls.

  • Concurrency and deduplication for crawlers/search: normalize URLs, use a thread-safe dedupe store (Bloom filter + persistent set), and enforce per-host rate limits; maintain politeness via robots.txt parsing.

Worked example — Design a REST API Abstraction Layer

First 30s: clarify constraints — client language surface (multiple SDKs?), sync vs async usage, SLOs for latency and availability, and which external failure modes must be masked. Skeleton pillars: (1) client contract and typed models (SDK), (2) reliability layer (timeouts, retries, circuit breaker, idempotency), (3) observability & tracing (structured logs, metrics, traces), (4) security & versioning (auth, API versioning). A concrete design decision: choose where to put retry logic — client SDK or server gateway — tradeoff being duplication across clients vs centralized control; mitigate duplicate side-effects by requiring idempotency keys. Explicitly call out performance tradeoffs: aggressive caching reduces latency but may serve stale data; aggressive retries improve success-rate but risk overloading slow third-parties. Close by saying what you'd do with more time: add adaptive retry policies driven by runtime metrics, add chaos-testing for dependency behaviors, and provide SDK code-gen for consistent contracts.

A second angle — Design resilient auth with flaky third-party tokens

Frame as an external dependency problem where auth token issuance/validation is brittle. Apply the same primitives: cache tokens and their expiry, proactively refresh with jittered retries, and run a circuit breaker around the token service so authentication degrades predictably (e.g., allow limited grace period using cached tokens). For verification, prefer locally verifiable tokens (signed JWT) to avoid synchronous validation calls; if not possible, use asynchronous validation with short-lived allowlists. Key tradeoff: shorter TTLs improve security but increase load on the token issuer; choose TTL to meet both security and availability SLOs and document the acceptable failure modes (e.g., read-only degraded access).

Common pitfalls

Pitfall: Unbounded retries and blocking threads.
Unbounded retries (or very long blocking timeouts) exhaust threadpools and create cascading failures. Size threadpools, use non-blocking IO where possible, cap retries, and apply circuit breakers.

Pitfall: Designing retries without idempotency.
Retrying mutating operations without idempotency leads to duplicated side-effects. Require idempotency keys or switch such operations to an at-most-once server workflow with dedupe stores.

Pitfall: Focusing only on mechanisms, not SLOs.
Listing patterns like "use cache or circuit breaker" without quantifying SLOs (99.9% success, p99 latency) leaves the interviewer unsure about tradeoffs. State target SLOs and show how each mechanism supports them.

Connections

Interviewers often pivot to adjacent topics: distributed tracing and correlation-id propagation for end-to-end debugging, rate-limiting & quota enforcement for multi-tenant systems, or the tradeoffs between sync vs async integration and replay semantics (exactly-once vs at-least-once).

Further reading

Practice questions

Related concepts