Ingest and Serve RAG Sessions with Up to 200 Documents
Company: Qualified Health
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
Design ingestion and serving for an enterprise Agent/RAG conversation system. A session can upload up to 200 documents, including PDF, DOCX, TXT, and CSV, totaling hundreds of megabytes. Users then ask questions across those documents and hold multi-turn conversations.
The reported targets are more than 100,000 active sessions, end-to-end response time below 3 seconds, and time to first streamed token below 500 milliseconds. Explain how you would interpret, budget, and validate these targets while preventing document processing and provider throttling from overwhelming the service.
### Constraints & Assumptions
- The source does not specify document page counts, parsing complexity, output length, request arrival rate, or provider quotas.
- Active sessions are not automatically simultaneous model generations.
- Upload, parsing, chunking, embedding, retrieval, and generation are distinct stages. Clarify which stages the response targets include.
- Documents and vectors expire after 24 hours in this source's broader requirements; serving and caches must respect that boundary.
### Clarifying Questions to Ask
- Do the query targets apply only after ingestion is complete, and does total response time include the final generated token?
- May queries use only the ready subset of documents, or must every document be searchable first?
- What is the distribution of concurrent queries and generated output lengths across active sessions?
- Which provider request/token quotas and regional latency constraints are available?
### Part 1 — Upload and Asynchronous Ingestion
Design the path from document upload to a searchable index without buffering or parsing every file in the web gateway. Explain status reporting, retries, and the indexing boundary visible to queries.
#### What This Part Should Cover
- Authorized bounded uploads and durable object references.
- Idempotent parsing/chunking/embedding stages and per-document readiness.
- A consistent corpus generation or declared partial-readiness policy.
### Part 2 — Meet and Measure Query Latency
Allocate the critical path and describe retrieval, context assembly, generation, and caching. Explain when semantic caching can safely reuse an answer and when it cannot.
#### What This Part Should Cover
- Separate TTFT and total-response budgets, with percentiles and output bounds clarified.
- Authorization-, corpus-, conversation-, and model-aware cache identity.
- Measurement of feasibility rather than a guarantee unsupported by provider or workload data.
### Part 3 — Scale Through Quota Pressure
Explain admission control, provider 429 handling, and backpressure across ingestion and interactive queries.
#### What This Part Should Cover
- Session concurrency versus active-request concurrency.
- Bounded request/token budgets, class separation, and fair admission.
- Retry delays and visible degradation without silently claiming the latency target was met.
```hint A ready session is not a ready corpus
The upload may be complete while parsing or embedding is still pending. Decide which state allows a query to promise coverage of all 200 documents.
```
### What a Strong Answer Covers
- A complete upload-to-index path and a query path with explicit readiness semantics.
- Honest feasibility reasoning for the reported latency and scale targets.
- Correctly scoped caches and stable behavior under model-provider rate limits.
### Follow-up Questions
- How should cached answers behave when a session's documents expire?
- Why can two semantically similar questions require different answers in different conversation turns?
- How would you keep a burst of background embeddings from consuming all interactive query quota?
Overview: Design a 200-document RAG session pipeline with asynchronous ingestion, corpus readiness, latency budgeting, scoped caching, and quota-aware serving.
Read the full Qualified Health Software Engineer interview experience this question came from