System design interview question: An Agent / RAG chat system supporting uploads of 200 documents
π Core Problem
Design an enterprise LLM agent question-answering system that lets users upload external documents in batches within one session, up to 200 documents, and have real-time streaming question-and-answer exchanges and multi-turn interactions about their contents.
π― System Requirements and Constraints (Requirements & SLAs)
- Document volume: Support uploads of up to 200 documents per session, including PDF, DOCX, TXT, CSV, and other formats, totaling up to hundreds of MB.
- Latency SLA: Total end-to-end response time must be < 3 seconds, and time to first streamed token (TTFT) must be < 500 milliseconds.
- Concurrency: Support 100,000+ concurrently active sessions.
- Data retention and compliance:
- Uploaded documents and vector indexes must automatically expire and be physically deleted after 24 hours, leaving no residual data.
- Users' historical conversation records must be retained for 7 days for review, and the sidebar list must load in under 10 milliseconds.
- Interruption: During generation, users must be able to click "Stop Generating" at any time and immediately terminate the model call.
β Interviewer's Core Follow-up Questions (Deep-Dive Questions)
1. Communication Protocols and Interrupting Streaming
- For LLM generation, should frontend/backend streaming use SSE (Server-Sent Events) or WebSocket? What are the advantages, disadvantages, and suitable scenarios for each?
- When a user clicks "Stop Generating," how should the frontend/backend network path be designed? How can downstream streaming stop within milliseconds while also actually interrupting the underlying model's generation and billing?
2. Ingesting and Chunking Large Document Collections (Ingestion Pipeline)
- With 200 documents totaling hundreds of MB uploaded in one session, how would you design ingestion and parsing without blocking the frontend UI or consuming large amounts of CPU and memory on the web gateway?
- For batches of large files, how should the asynchronous task flow for text parsing, chunking, and embeddings be orchestrated? How should its status be synchronized to the frontend?
3. Fast Responses and Meeting Latency Targets (Latency SLA)
- With an extremely long knowledge base built from 200 uploaded documents, how can end-to-end latency stay under 3 seconds and time to first token under 500 milliseconds?
- How should caching be designed in this scenario? Can semantically similar questions receive direct responses within milliseconds?
4. Managing Two Data Lifecycles (Data Retention & Privacy)
- Why should documents and vectors expire after 24 hours while chat records are retained for 7 days? How should these two completely different retention policies be implemented in underlying storage?
- Vector databases generally don't support a per-document TTL on individual entries. How can vectors be physically cleaned up automatically as soon as 24 hours have elapsed, avoiding storage growth and ensuring compliance?
- When a user opens a historical conversation on day 4, the underlying documents have already expired. How should the user experience and product state machine work? What should happen if the user continues asking questions?
5. Historical Retrieval Performance Across Sessions
- How can the sidebar load all conversation history from the past 7 days with an extremely low latency of under 10 milliseconds? How should the underlying database and indexes be designed?
6. Multi-Tenant Data Isolation (Multi-Tenancy)
- When many users share the same underlying vector-database cluster and indexes, how should the architecture strictly prevent unauthorized data access and data leaks across tenants and sessions?
7. Defense in Depth and Safety Guardrails (Security & Guardrails)
- Should security controls, such as prompt-injection defenses, sensitive-data redaction, jailbreak prevention, and rate limiting, sit at the traditional web-gateway layer or the model-guardrail layer? What are their respective responsibilities and boundaries?
8. High Concurrency and Resilience to Model Rate Limits (Scaling & Throttling)
- If a sudden traffic surge causes the underlying LLM API to frequently return HTTP 429, due to throttling or reaching the quota limit, what architectural approaches can provide resilience and smooth traffic peaks?
9. Failures Along the Call Chain and Retries (Fault Tolerance)
- In a complex call chain involving agent orchestration, a foundation model, a vector-retrieval database, and external tools/actions, how should the system respond if any node times out or crashes?
- Why can't an agent architecture blindly retry everything automatically? How can external tool calls avoid duplicate side effects, such as duplicate charges or duplicate database writes?
10. Long-Term Compliance Archives (Cold Storage & Audit)
- If an enterprise requires read-only audit access to conversation logs for months or even years, how can you design separate hot and cold storage and log analysis without slowing the production database or causing storage costs to explode?
Discussion
Loading commentsβ¦