Design input validation and error handling

Quick Overview

A Meta software engineer system design technical screen: design the API and system-level input validation and error-handling strategy for a service that computes the top-k most frequent elements in a posted array. The merged answer covers the request/response contract with a worked example, payload and cardinality limits, layered validation across client, gateway and service, authentication, rate limiting, idempotency, timeouts and retries, the full 4xx/5xx error taxonomy with sample error responses, partial-failure modes, graceful degradation to approximate top-k with stated error bounds, defenses against compression bombs and poisoned payloads, observability, and a test and production monitoring plan.

Design input validation and error handling

Company: Meta

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Technical Screen

##### Question You own a service endpoint that computes the **top-k most frequent elements** in a posted array. The counting algorithm is not the interview — the contract around it is. Design the end-to-end API and system-level input-validation and error-handling strategy for that endpoint. Walk through: 1. **API contract** — endpoint, headers, request and response schema. Give a concrete example request and a concrete success response. 2. **Limits and constraints** — payload size, array length, `k` bounds, allowed element types and element size, compression. 3. **Validation layers** — what the client, the gateway/WAF, and the service each enforce, in what order, and how the schema is enforced and versioned. 4. **Authentication and authorization** — external clients vs service-to-service, tenant isolation, scopes. 5. **Rate limiting and quotas** — what you meter, and what a client sees when it trips a limit. 6. **Idempotency** — how a client retries a POST safely, and what happens when a key is reused with a different body. 7. **Timeouts and retries** — the per-phase time budget, which status you return when it is exceeded, and the retry/backoff policy you hand clients. 8. **Error taxonomy** — client (4xx) vs server (5xx) errors, stable machine-readable codes, and one consistent error envelope. Show sample error responses. 9. **Partial-failure behavior** — what happens when *some* elements are invalid: reject the whole request, or skip them and report? 10. **Graceful degradation** — behavior for oversized but well-formed payloads: exact vs approximate top-k, sampling, load shedding. 11. **Safeguards against malformed or poisoned payloads** — compression bombs, invalid UTF-8, deep nesting, duplicate JSON keys, adversarially high cardinality. 12. **Observability** — structured logs, metrics, and traces, and what must never appear in any of them. 13. **Testing and production monitoring** — how you would test this before ship, and what you would alert on after. Be explicit about the defaults you pick and the trade-off behind each one. The exact numbers matter less than whether you can defend them.

Overview: A Meta software engineer system design technical screen: design the API and system-level input validation and error-handling strategy for a service that computes the top-k most frequent elements in a posted array. The merged answer covers the request/response contract with a worked example, payload and cardinality limits, layered validation across client, gateway and service, authentication, rate limiting, idempotency, timeouts and retries, the full 4xx/5xx error taxonomy with sample error responses, partial-failure modes, graceful degradation to approximate top-k with stated error bounds, defenses against compression bombs and poisoned payloads, observability, and a test and production monitoring plan.

|Home/System Design/Meta
Meta logo
Meta
Sep 6, 2025
hardSoftware EngineerTechnical ScreenSystem Design
6
0
Question

You own a service endpoint that computes the top-k most frequent elements in a posted array. The counting algorithm is not the interview — the contract around it is. Design the end-to-end API and system-level input-validation and error-handling strategy for that endpoint.

Walk through:

  1. API contract — endpoint, headers, request and response schema. Give a concrete example request and a concrete success response.
  2. Limits and constraints — payload size, array length, k bounds, allowed element types and element size, compression.
  3. Validation layers — what the client, the gateway/WAF, and the service each enforce, in what order, and how the schema is enforced and versioned.
  4. Authentication and authorization — external clients vs service-to-service, tenant isolation, scopes.
  5. Rate limiting and quotas — what you meter, and what a client sees when it trips a limit.
  6. Idempotency — how a client retries a POST safely, and what happens when a key is reused with a different body.
  7. Timeouts and retries — the per-phase time budget, which status you return when it is exceeded, and the retry/backoff policy you hand clients.
  8. Error taxonomy — client (4xx) vs server (5xx) errors, stable machine-readable codes, and one consistent error envelope. Show sample error responses.
  9. Partial-failure behavior — what happens when some elements are invalid: reject the whole request, or skip them and report?
  10. Graceful degradation — behavior for oversized but well-formed payloads: exact vs approximate top-k, sampling, load shedding.
  11. Safeguards against malformed or poisoned payloads — compression bombs, invalid UTF-8, deep nesting, duplicate JSON keys, adversarially high cardinality.
  12. Observability — structured logs, metrics, and traces, and what must never appear in any of them.
  13. Testing and production monitoring — how you would test this before ship, and what you would alert on after.

Be explicit about the defaults you pick and the trade-off behind each one. The exact numbers matter less than whether you can defend them.

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...