Ingest and Serve RAG Sessions with Up to 200 Documents

Read the full interview experience this question came from →

Quick Overview

Design a 200-document RAG session pipeline with asynchronous ingestion, corpus readiness, latency budgeting, scoped caching, and quota-aware serving.

Ingest and Serve RAG Sessions with Up to 200 Documents

Company: Qualified Health

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

Design ingestion and serving for an enterprise Agent/RAG conversation system. A session can upload up to 200 documents, including PDF, DOCX, TXT, and CSV, totaling hundreds of megabytes. Users then ask questions across those documents and hold multi-turn conversations. The reported targets are more than 100,000 active sessions, end-to-end response time below 3 seconds, and time to first streamed token below 500 milliseconds. Explain how you would interpret, budget, and validate these targets while preventing document processing and provider throttling from overwhelming the service. ### Constraints & Assumptions - The source does not specify document page counts, parsing complexity, output length, request arrival rate, or provider quotas. - Active sessions are not automatically simultaneous model generations. - Upload, parsing, chunking, embedding, retrieval, and generation are distinct stages. Clarify which stages the response targets include. - Documents and vectors expire after 24 hours in this source's broader requirements; serving and caches must respect that boundary. ### Clarifying Questions to Ask - Do the query targets apply only after ingestion is complete, and does total response time include the final generated token? - May queries use only the ready subset of documents, or must every document be searchable first? - What is the distribution of concurrent queries and generated output lengths across active sessions? - Which provider request/token quotas and regional latency constraints are available? ### Part 1 — Upload and Asynchronous Ingestion Design the path from document upload to a searchable index without buffering or parsing every file in the web gateway. Explain status reporting, retries, and the indexing boundary visible to queries. #### What This Part Should Cover - Authorized bounded uploads and durable object references. - Idempotent parsing/chunking/embedding stages and per-document readiness. - A consistent corpus generation or declared partial-readiness policy. ### Part 2 — Meet and Measure Query Latency Allocate the critical path and describe retrieval, context assembly, generation, and caching. Explain when semantic caching can safely reuse an answer and when it cannot. #### What This Part Should Cover - Separate TTFT and total-response budgets, with percentiles and output bounds clarified. - Authorization-, corpus-, conversation-, and model-aware cache identity. - Measurement of feasibility rather than a guarantee unsupported by provider or workload data. ### Part 3 — Scale Through Quota Pressure Explain admission control, provider 429 handling, and backpressure across ingestion and interactive queries. #### What This Part Should Cover - Session concurrency versus active-request concurrency. - Bounded request/token budgets, class separation, and fair admission. - Retry delays and visible degradation without silently claiming the latency target was met. ```hint A ready session is not a ready corpus The upload may be complete while parsing or embedding is still pending. Decide which state allows a query to promise coverage of all 200 documents. ``` ### What a Strong Answer Covers - A complete upload-to-index path and a query path with explicit readiness semantics. - Honest feasibility reasoning for the reported latency and scale targets. - Correctly scoped caches and stable behavior under model-provider rate limits. ### Follow-up Questions - How should cached answers behave when a session's documents expire? - Why can two semantically similar questions require different answers in different conversation turns? - How would you keep a burst of background embeddings from consuming all interactive query quota?

Overview: Design a 200-document RAG session pipeline with asynchronous ingestion, corpus readiness, latency budgeting, scoped caching, and quota-aware serving.

Read the full Qualified Health Software Engineer interview experience this question came from

|Home/System Design/Qualified Health
Qualified Health logo
Qualified Health
Sep 8, 2026
hardSoftware EngineerOnsiteSystem Design
0
0

Design ingestion and serving for an enterprise Agent/RAG conversation system. A session can upload up to 200 documents, including PDF, DOCX, TXT, and CSV, totaling hundreds of megabytes. Users then ask questions across those documents and hold multi-turn conversations.

The reported targets are more than 100,000 active sessions, end-to-end response time below 3 seconds, and time to first streamed token below 500 milliseconds. Explain how you would interpret, budget, and validate these targets while preventing document processing and provider throttling from overwhelming the service.

Constraints & Assumptions

  • The source does not specify document page counts, parsing complexity, output length, request arrival rate, or provider quotas.
  • Active sessions are not automatically simultaneous model generations.
  • Upload, parsing, chunking, embedding, retrieval, and generation are distinct stages. Clarify which stages the response targets include.
  • Documents and vectors expire after 24 hours in this source's broader requirements; serving and caches must respect that boundary.

Clarifying Questions to Ask Guidance

  • Do the query targets apply only after ingestion is complete, and does total response time include the final generated token?
  • May queries use only the ready subset of documents, or must every document be searchable first?
  • What is the distribution of concurrent queries and generated output lengths across active sessions?
  • Which provider request/token quotas and regional latency constraints are available?

Part 1 — Upload and Asynchronous Ingestion

Design the path from document upload to a searchable index without buffering or parsing every file in the web gateway. Explain status reporting, retries, and the indexing boundary visible to queries.

What This Part Should Cover Guidance

  • Authorized bounded uploads and durable object references.
  • Idempotent parsing/chunking/embedding stages and per-document readiness.
  • A consistent corpus generation or declared partial-readiness policy.

Part 2 — Meet and Measure Query Latency

Allocate the critical path and describe retrieval, context assembly, generation, and caching. Explain when semantic caching can safely reuse an answer and when it cannot.

What This Part Should Cover Guidance

  • Separate TTFT and total-response budgets, with percentiles and output bounds clarified.
  • Authorization-, corpus-, conversation-, and model-aware cache identity.
  • Measurement of feasibility rather than a guarantee unsupported by provider or workload data.

Part 3 — Scale Through Quota Pressure

Explain admission control, provider 429 handling, and backpressure across ingestion and interactive queries.

What This Part Should Cover Guidance

  • Session concurrency versus active-request concurrency.
  • Bounded request/token budgets, class separation, and fair admission.
  • Retry delays and visible degradation without silently claiming the latency target was met.

What a Strong Answer Covers Guidance

  • A complete upload-to-index path and a query path with explicit readiness semantics.
  • Honest feasibility reasoning for the reported latency and scale targets.
  • Correctly scoped caches and stable behavior under model-provider rate limits.

Follow-up Questions Guidance

  • How should cached answers behave when a session's documents expire?
  • Why can two semantically similar questions require different answers in different conversation turns?
  • How would you keep a burst of background embeddings from consuming all interactive query quota?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...