Design and optimize a RAG system
Company: OpenAI
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Onsite
##### Question
You are building a **Retrieval-Augmented Generation (RAG)** system for question answering over a private corpus — engineering wikis, design docs, runbooks, support tickets, PDFs, knowledge-base articles. Users ask natural-language questions in an interactive chat surface and expect grounded answers **with citations** back to the source documents.
Design the **end-to-end system**, describe the **optimizations** that make it high-quality, low-latency and trustworthy in production, and — for every component — explain **how you would evaluate it**. The evaluation half is not an afterthought here: for each stage the interviewer wants a named failure mode, the metric that detects it, and how you would get the labels.
Work through the following parts:
1. **End-to-end architecture and data flow** — the full component set (ingestion, parsing, chunking, embeddings, indexing, retrieval, reranking, generation) and the split between the asynchronous indexing path and the synchronous query path.
2. **Ingestion and preprocessing** — connectors, change detection, and layout-aware parsing of heterogeneous formats (PDF, HTML, wiki, Markdown), including tables and scanned documents.
3. **Chunking** — splitting strategy for long, structured documents; size and overlap trade-offs; what metadata rides along with each chunk.
4. **Embeddings and indexing** — embedding choice and versioning, ANN index, a parallel lexical (BM25) index, and metadata/ACL-filterable fields.
5. **Retrieval** — query understanding and rewriting, top-N, metadata filters, hybrid sparse+dense search and fusion.
6. **Reranking** — whether to add a second stage, what it buys, and what it costs.
7. **Context assembly and generation** — prompt construction, dedupe and ordering, context compression, citation format.
8. **Grounding, guardrails and fallback** — how you reduce hallucinations, and exactly what the system does when retrieval comes back weak or the question is ambiguous.
9. **Continuous ingestion and freshness** — inserts, **edits**, and **deletions**, and the guarantee that deleted or superseded content can never be retrieved or cited.
10. **Access control** — enforcing per-user ACLs without leaking even the existence of a forbidden document.
11. **Latency and cost** — how you hit the interactive latency budget and where you trade quality for latency or cost.
12. **Evaluation plan, stage by stage** — ingestion/chunking quality, retrieval quality, reranking quality, generation quality and grounding, and end-to-end user success. Cover both offline and online evaluation, and say **how you obtain labels** if no gold set exists.
13. **Monitoring, drift and the improvement loop** — what you log, what you alert on, how you detect corpus/query drift, and how production failures become regression tests.
### Constraints & assumptions
State your own numbers explicitly, but a reasonable baseline to design against:
- Corpus on the order of **10^6–10^7 chunks** after splitting, heterogeneous formats, long documents.
- Interactive latency target: **p95 end-to-end < 2 s** to first token, with answer streaming.
- Ingestion freshness SLA: a new, edited or deleted document is reflected in retrieval within **minutes**, not hours.
- Documents carry **access-control metadata (ACLs)** — not every user may see every document.
- Citations are **mandatory**; abstention ("I couldn't find this in the docs") is allowed.
- You may assume access to a managed or self-hosted **vector index**, a **lexical/BM25 index**, an embedding model, a cross-encoder reranker, and a generation LLM.
### Clarifying questions to ask
A strong candidate scopes the problem before designing:
- What is the **expected query volume** (QPS) and the read/write ratio (query rate vs. ingestion rate)?
- What is the **query mix** — single-fact lookup, multi-hop reasoning, summarization, or "list all X"? This decides chunk size, top-k, and whether you need iterative retrieval.
- Are answers **single-turn or conversational** (must the retriever resolve follow-up references against chat history)?
- How strict are the **access-control and data-isolation** requirements — must we prevent even leaking the *existence* of a forbidden document?
- What is the **cost budget** per query (embedding calls, reranker calls, generation tokens)?
- How will quality be judged — is there an existing **labeled eval set**, or do we need to bootstrap one?
- What is the **output contract** — free text or structured JSON, and what is the tolerance for abstention versus always producing something?
```hint Where to start
Frame it as two decoupled planes: an **offline/asynchronous indexing plane** (ingest → parse → chunk → embed → index) and an **online query plane** (query understanding → retrieve → rerank → generate → verify). Sketch both separately so the freshness budget lives in one and the latency budget lives in the other; they meet only at the shared indices.
```
```hint Retrieval quality
The highest-leverage relevance move is usually **hybrid retrieval** (sparse lexical like BM25 + dense vector) fused with Reciprocal Rank Fusion, followed by a **cross-encoder reranker** over the top-N candidates. Think about why pure dense retrieval fails on exact tokens — error codes, IDs, version strings, function names.
```
```hint Grounding
Reducing hallucination is a chain, not one trick: **grounded prompting** ("answer only from the provided sources"), **citation enforcement**, **confidence gating** (abstain when the top scores are low), and an optional **faithfulness/entailment check** of the generated answer against the retrieved passages.
```
```hint Evaluating each stage
For every component, answer three questions: what goes wrong here, which metric would show it, and where do the labels come from? The one candidates usually skip is ingestion/chunking — a good proxy is what fraction of gold answer spans fall entirely inside a single chunk.
```
### Follow-up questions
- A user reports the system **confidently cited the wrong document**. Walk through how you localize the failure — chunk never retrieved, retrieved but mis-ranked, or present in the prompt and ignored — and what you change for each.
- Your **embedding model is upgraded** to a new version. What is your re-indexing and rollout plan, and how do you avoid a mixed-embedding-space index where old and new vectors are incomparable?
- How would you extend the system to handle **multi-hop questions**, and how does that differ from **"list all / aggregate"** questions?
- The corpus contains **near-duplicate documents** (e.g. copies of the same runbook). How does this hurt retrieval, and how would you handle it?
- Two retrieved documents **contradict each other**. What should the system do?
- You use an **LLM as a judge** in your offline eval. How do you know the judge is any good, and what would make you stop trusting it?
Overview: An OpenAI machine learning engineer onsite ML system design question: design, optimize and evaluate a production Retrieval-Augmented Generation (RAG) system that answers questions over a private corpus with citations. The merged answer covers the full pipeline — ingestion and layout-aware parsing, structure-aware chunking, embeddings and versioning, hybrid BM25 plus vector retrieval with RRF fusion, cross-encoder reranking, grounded generation with forced citations, abstention and a fallback ladder, ACL enforcement during retrieval, and the insert/edit/delete freshness lifecycle. It then gives a stage-by-stage evaluation plan (chunk answerability, hit-rate@K, recall@K, nDCG, faithfulness, citation quality, end-to-end success), how to bootstrap labels with no gold set, and follow-ups on debugging a wrongly cited document, embedding-model upgrades, multi-hop versus list-all queries, near-duplicates, conflicting sources and validating an LLM judge.
Read the full OpenAI Machine Learning Engineer interview experience this question came from
Community answers
Answer by Jack
chunking is also important to cover