Design input validation and error handling
Company: Meta
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Technical Screen
##### Question
You own a service endpoint that computes the **top-k most frequent elements** in a posted array. The counting algorithm is not the interview — the contract around it is. Design the end-to-end API and system-level input-validation and error-handling strategy for that endpoint.
Walk through:
1. **API contract** — endpoint, headers, request and response schema. Give a concrete example request and a concrete success response.
2. **Limits and constraints** — payload size, array length, `k` bounds, allowed element types and element size, compression.
3. **Validation layers** — what the client, the gateway/WAF, and the service each enforce, in what order, and how the schema is enforced and versioned.
4. **Authentication and authorization** — external clients vs service-to-service, tenant isolation, scopes.
5. **Rate limiting and quotas** — what you meter, and what a client sees when it trips a limit.
6. **Idempotency** — how a client retries a POST safely, and what happens when a key is reused with a different body.
7. **Timeouts and retries** — the per-phase time budget, which status you return when it is exceeded, and the retry/backoff policy you hand clients.
8. **Error taxonomy** — client (4xx) vs server (5xx) errors, stable machine-readable codes, and one consistent error envelope. Show sample error responses.
9. **Partial-failure behavior** — what happens when *some* elements are invalid: reject the whole request, or skip them and report?
10. **Graceful degradation** — behavior for oversized but well-formed payloads: exact vs approximate top-k, sampling, load shedding.
11. **Safeguards against malformed or poisoned payloads** — compression bombs, invalid UTF-8, deep nesting, duplicate JSON keys, adversarially high cardinality.
12. **Observability** — structured logs, metrics, and traces, and what must never appear in any of them.
13. **Testing and production monitoring** — how you would test this before ship, and what you would alert on after.
Be explicit about the defaults you pick and the trade-off behind each one. The exact numbers matter less than whether you can defend them.
Overview: A Meta software engineer system design technical screen: design the API and system-level input validation and error-handling strategy for a service that computes the top-k most frequent elements in a posted array. The merged answer covers the request/response contract with a worked example, payload and cardinality limits, layered validation across client, gateway and service, authentication, rate limiting, idempotency, timeouts and retries, the full 4xx/5xx error taxonomy with sample error responses, partial-failure modes, graceful degradation to approximate top-k with stated error bounds, defenses against compression bombs and poisoned payloads, observability, and a test and production monitoring plan.