Resilient APIs And External Dependency Integration
Asked of: Software Engineer
Last updated
What's being tested
Interviewers probe your ability to integrate and expose external services reliably: designing APIs and clients that tolerate flakiness, control resource usage, and fail predictably. Expect questions that check knowledge of timeouts, retries, idempotency, backpressure and isolation (bulkheads), and observable failure modes (p99 latency, success rate). Snowflake values engineers who can keep user-facing paths available and debuggable while respecting external SLA variability.
Core knowledge
-
Retries: implement exponential backoff with jitter (e.g., base b, attempts n → delay ≈ b * 2^n ± jitter) and a capped max attempts to avoid cascading load.
-
Idempotency: require idempotency keys for non-idempotent server actions; client SDK should attach
Idempotency-Keyand deduplicate on the server to avoid side-effect duplication. -
Circuit breaker: use circuit breaker (closed/half-open/open) to short-circuit calls after threshold failures; combine with health checks and cooldown windows to prevent retry storms.
-
Bulkhead isolation: partition pools by external dependency or operation type (threadpool/connection pool) so one slow API cannot exhaust all resources; size pools using capacity planning.
-
Timeouts vs deadlines: use per-call timeout shorter than user-facing deadline, and propagate a deadline header so downstreams can fail fast; avoid infinite blocking.
-
Rate limiting: enforce
token bucketorleaky bucketper-tenant and per-external-host; implement global and per-host quotas for crawlers or outward API calls. -
Caching & staleness: use read-through caches with TTL and
stale-while-revalidatefor soft degradation; consider cache invalidation costs — LRU works up to ~10M keys; beyond that use approximate structures or sharding. -
Backpressure & queues: buffer with bounded queues; apply admission control (reject or degrade) when queue length > threshold; Little’s law applies: to size pools.
-
Observability: emit structured logs, metrics (
request_rate,error_rate,p50/p95/p99), and distributed traces with a correlation id; alert on success-rate SLOs, not just latency. -
Async fallback & DLQ: convert sync calls to async via queues for brittle dependencies; failed items go to a dead-letter queue for investigation and replay.
-
Security & token handling: cache third-party tokens, refresh proactively before expiry; on flaky token endpoints, use retry+Jitter and local validation (e.g., verify signatures) to reduce calls.
-
Concurrency and deduplication for crawlers/search: normalize URLs, use a thread-safe dedupe store (Bloom filter + persistent set), and enforce per-host rate limits; maintain politeness via
robots.txtparsing.
Worked example — Design a REST API Abstraction Layer
First 30s: clarify constraints — client language surface (multiple SDKs?), sync vs async usage, SLOs for latency and availability, and which external failure modes must be masked. Skeleton pillars: (1) client contract and typed models (SDK), (2) reliability layer (timeouts, retries, circuit breaker, idempotency), (3) observability & tracing (structured logs, metrics, traces), (4) security & versioning (auth, API versioning). A concrete design decision: choose where to put retry logic — client SDK or server gateway — tradeoff being duplication across clients vs centralized control; mitigate duplicate side-effects by requiring idempotency keys. Explicitly call out performance tradeoffs: aggressive caching reduces latency but may serve stale data; aggressive retries improve success-rate but risk overloading slow third-parties. Close by saying what you'd do with more time: add adaptive retry policies driven by runtime metrics, add chaos-testing for dependency behaviors, and provide SDK code-gen for consistent contracts.
A second angle — Design resilient auth with flaky third-party tokens
Frame as an external dependency problem where auth token issuance/validation is brittle. Apply the same primitives: cache tokens and their expiry, proactively refresh with jittered retries, and run a circuit breaker around the token service so authentication degrades predictably (e.g., allow limited grace period using cached tokens). For verification, prefer locally verifiable tokens (signed JWT) to avoid synchronous validation calls; if not possible, use asynchronous validation with short-lived allowlists. Key tradeoff: shorter TTLs improve security but increase load on the token issuer; choose TTL to meet both security and availability SLOs and document the acceptable failure modes (e.g., read-only degraded access).
Common pitfalls
Pitfall: Unbounded retries and blocking threads.
Unbounded retries (or very long blocking timeouts) exhaust threadpools and create cascading failures. Size threadpools, use non-blocking IO where possible, cap retries, and apply circuit breakers.
Pitfall: Designing retries without idempotency.
Retrying mutating operations without idempotency leads to duplicated side-effects. Require idempotency keys or switch such operations to an at-most-once server workflow with dedupe stores.
Pitfall: Focusing only on mechanisms, not SLOs.
Listing patterns like "use cache or circuit breaker" without quantifying SLOs (99.9%success,p99latency) leaves the interviewer unsure about tradeoffs. State target SLOs and show how each mechanism supports them.
Connections
Interviewers often pivot to adjacent topics: distributed tracing and correlation-id propagation for end-to-end debugging, rate-limiting & quota enforcement for multi-tenant systems, or the tradeoffs between sync vs async integration and replay semantics (exactly-once vs at-least-once).
Further reading
-
Release It! by Michael T. Nygard — practical resilience patterns, circuit breakers and bulkheads.
-
Martin Fowler — Circuit Breaker — concise explanation of the pattern and common thresholds.
Practice questions
- Find the Shortest Click Path with a Fallible Link APISnowflake · Software Engineer · Technical Screen · medium
- Design an Interactive Query Execution NotebookSnowflake · Software Engineer · Technical Screen · medium
- Design a REST API Abstraction LayerSnowflake · Software Engineer · Technical Screen · hard
- Design a concurrent web crawlerSnowflake · Software Engineer · Onsite · hard
- Design resilient auth with flaky third-party tokensSnowflake · Software Engineer · Onsite · hard
Related concepts
- Resilient API Aggregation And Operational DebuggingSoftware Engineering Fundamentals
- API Integration And External Service DesignSystem Design
- Idempotent API DesignSystem Design
- Async Services, Idempotency, And RetriesSystem Design
- Distributed Systems Reliability And StorageSystem Design
- API Idempotency And Concurrency ControlSystem Design