Control LLM Workflow Quotas and Recovery Costs
Company: Qualified Health
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
A distributed DAG workflow system calls external LLM providers for document summarization, embeddings, extraction, and inference. Model tasks may take 10–60 seconds, process long documents, and fail after part of their work has completed. The reported aggregate demand is about 80 billion tokens per day, subject to provider request-per-minute and token-per-minute quotas.
Design quota-aware dispatch and recovery that handle HTTP 429 responses without overwhelming the cluster and avoid unnecessary repeat model work after a late failure.
### Constraints & Assumptions
- The reported token demand is a sizing input, not proof that a provider grants that capacity.
- Quota scopes, headers, idempotency, asynchronous job lookup, cancellation, and usage-reporting capabilities differ by provider; do not assume unsupported features.
- A transport timeout does not establish whether an external request ran or was billed.
- Document processing may be decomposable into chunks, but preserve any semantic dependencies among chunk outputs and the final result.
### Clarifying Questions to Ask
- Which quota applies to each provider, account, model, region, request, and token class?
- Can token usage be estimated before dispatch and reconciled with actual usage afterward?
- Can accepted requests be queried or retried with the same idempotency key?
- Which intermediate document results are reusable under the same input, model, prompt, and transformation versions?
### Part 1 — Quota Admission and Backpressure
Design distributed quota accounting, concurrency limits, delayed retries, and tenant fairness. Explain how a surge of HTTP 429 responses changes admission instead of creating a synchronized retry storm.
#### What This Part Should Cover
- Separate request and token budgets, including in-flight reservations.
- Shared quota ownership across workers and bounded local allowances if used.
- Backoff, jitter, retry limits, and propagation of pressure to upstream workflow admission.
### Part 2 — Checkpoint Useful Work and Reconcile Uncertain Calls
A long-document task fails near its end. Explain what is durably retained, which work can safely resume, and what remains uncertain if the last provider response was lost.
#### What This Part Should Cover
- Version-bound chunk outputs and completion records.
- Stable operation identity and provider-dependent reconciliation.
- The limits of guarantees about duplicate generation and billing.
```hint Quota belongs to a scope
Several workers may share one provider account. Giving every worker the full account quota multiplies the allowed traffic instead of coordinating it.
```
### What a Strong Answer Covers
- A bounded quota-aware dispatch path that remains stable during rate limiting.
- Recovery at meaningful document-processing boundaries.
- Separation of confirmed reusable results from possibly completed or billed calls with missing responses.
### Follow-up Questions
- What should happen when a prompt or model version changes after some chunks have completed?
- How would you prevent one tenant's retries from consuming all newly available token capacity?
- Which guarantee is possible when a provider offers neither request lookup nor idempotent retry?
Overview: Control LLM workflow quotas and 429 retries with shared token budgets, durable checkpoints, and explicit handling of uncertain requests and duplicate billing.
Read the full Qualified Health Software Engineer interview experience this question came from