Design and optimize a RAG system

Read the full interview experience this question came from →

Quick Overview

An OpenAI machine learning engineer onsite ML system design question: design, optimize and evaluate a production Retrieval-Augmented Generation (RAG) system that answers questions over a private corpus with citations. The merged answer covers the full pipeline — ingestion and layout-aware parsing, structure-aware chunking, embeddings and versioning, hybrid BM25 plus vector retrieval with RRF fusion, cross-encoder reranking, grounded generation with forced citations, abstention and a fallback ladder, ACL enforcement during retrieval, and the insert/edit/delete freshness lifecycle. It then gives a stage-by-stage evaluation plan (chunk answerability, hit-rate@K, recall@K, nDCG, faithfulness, citation quality, end-to-end success), how to bootstrap labels with no gold set, and follow-ups on debugging a wrongly cited document, embedding-model upgrades, multi-hop versus list-all queries, near-duplicates, conflicting sources and validating an LLM judge.

Design and optimize a RAG system

Company: OpenAI

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: hard

Interview Round: Onsite

##### Question You are building a **Retrieval-Augmented Generation (RAG)** system for question answering over a private corpus — engineering wikis, design docs, runbooks, support tickets, PDFs, knowledge-base articles. Users ask natural-language questions in an interactive chat surface and expect grounded answers **with citations** back to the source documents. Design the **end-to-end system**, describe the **optimizations** that make it high-quality, low-latency and trustworthy in production, and — for every component — explain **how you would evaluate it**. The evaluation half is not an afterthought here: for each stage the interviewer wants a named failure mode, the metric that detects it, and how you would get the labels. Work through the following parts: 1. **End-to-end architecture and data flow** — the full component set (ingestion, parsing, chunking, embeddings, indexing, retrieval, reranking, generation) and the split between the asynchronous indexing path and the synchronous query path. 2. **Ingestion and preprocessing** — connectors, change detection, and layout-aware parsing of heterogeneous formats (PDF, HTML, wiki, Markdown), including tables and scanned documents. 3. **Chunking** — splitting strategy for long, structured documents; size and overlap trade-offs; what metadata rides along with each chunk. 4. **Embeddings and indexing** — embedding choice and versioning, ANN index, a parallel lexical (BM25) index, and metadata/ACL-filterable fields. 5. **Retrieval** — query understanding and rewriting, top-N, metadata filters, hybrid sparse+dense search and fusion. 6. **Reranking** — whether to add a second stage, what it buys, and what it costs. 7. **Context assembly and generation** — prompt construction, dedupe and ordering, context compression, citation format. 8. **Grounding, guardrails and fallback** — how you reduce hallucinations, and exactly what the system does when retrieval comes back weak or the question is ambiguous. 9. **Continuous ingestion and freshness** — inserts, **edits**, and **deletions**, and the guarantee that deleted or superseded content can never be retrieved or cited. 10. **Access control** — enforcing per-user ACLs without leaking even the existence of a forbidden document. 11. **Latency and cost** — how you hit the interactive latency budget and where you trade quality for latency or cost. 12. **Evaluation plan, stage by stage** — ingestion/chunking quality, retrieval quality, reranking quality, generation quality and grounding, and end-to-end user success. Cover both offline and online evaluation, and say **how you obtain labels** if no gold set exists. 13. **Monitoring, drift and the improvement loop** — what you log, what you alert on, how you detect corpus/query drift, and how production failures become regression tests. ### Constraints & assumptions State your own numbers explicitly, but a reasonable baseline to design against: - Corpus on the order of **10^6–10^7 chunks** after splitting, heterogeneous formats, long documents. - Interactive latency target: **p95 end-to-end < 2 s** to first token, with answer streaming. - Ingestion freshness SLA: a new, edited or deleted document is reflected in retrieval within **minutes**, not hours. - Documents carry **access-control metadata (ACLs)** — not every user may see every document. - Citations are **mandatory**; abstention ("I couldn't find this in the docs") is allowed. - You may assume access to a managed or self-hosted **vector index**, a **lexical/BM25 index**, an embedding model, a cross-encoder reranker, and a generation LLM. ### Clarifying questions to ask A strong candidate scopes the problem before designing: - What is the **expected query volume** (QPS) and the read/write ratio (query rate vs. ingestion rate)? - What is the **query mix** — single-fact lookup, multi-hop reasoning, summarization, or "list all X"? This decides chunk size, top-k, and whether you need iterative retrieval. - Are answers **single-turn or conversational** (must the retriever resolve follow-up references against chat history)? - How strict are the **access-control and data-isolation** requirements — must we prevent even leaking the *existence* of a forbidden document? - What is the **cost budget** per query (embedding calls, reranker calls, generation tokens)? - How will quality be judged — is there an existing **labeled eval set**, or do we need to bootstrap one? - What is the **output contract** — free text or structured JSON, and what is the tolerance for abstention versus always producing something? ```hint Where to start Frame it as two decoupled planes: an **offline/asynchronous indexing plane** (ingest → parse → chunk → embed → index) and an **online query plane** (query understanding → retrieve → rerank → generate → verify). Sketch both separately so the freshness budget lives in one and the latency budget lives in the other; they meet only at the shared indices. ``` ```hint Retrieval quality The highest-leverage relevance move is usually **hybrid retrieval** (sparse lexical like BM25 + dense vector) fused with Reciprocal Rank Fusion, followed by a **cross-encoder reranker** over the top-N candidates. Think about why pure dense retrieval fails on exact tokens — error codes, IDs, version strings, function names. ``` ```hint Grounding Reducing hallucination is a chain, not one trick: **grounded prompting** ("answer only from the provided sources"), **citation enforcement**, **confidence gating** (abstain when the top scores are low), and an optional **faithfulness/entailment check** of the generated answer against the retrieved passages. ``` ```hint Evaluating each stage For every component, answer three questions: what goes wrong here, which metric would show it, and where do the labels come from? The one candidates usually skip is ingestion/chunking — a good proxy is what fraction of gold answer spans fall entirely inside a single chunk. ``` ### Follow-up questions - A user reports the system **confidently cited the wrong document**. Walk through how you localize the failure — chunk never retrieved, retrieved but mis-ranked, or present in the prompt and ignored — and what you change for each. - Your **embedding model is upgraded** to a new version. What is your re-indexing and rollout plan, and how do you avoid a mixed-embedding-space index where old and new vectors are incomparable? - How would you extend the system to handle **multi-hop questions**, and how does that differ from **"list all / aggregate"** questions? - The corpus contains **near-duplicate documents** (e.g. copies of the same runbook). How does this hurt retrieval, and how would you handle it? - Two retrieved documents **contradict each other**. What should the system do? - You use an **LLM as a judge** in your offline eval. How do you know the judge is any good, and what would make you stop trusting it?

Overview: An OpenAI machine learning engineer onsite ML system design question: design, optimize and evaluate a production Retrieval-Augmented Generation (RAG) system that answers questions over a private corpus with citations. The merged answer covers the full pipeline — ingestion and layout-aware parsing, structure-aware chunking, embeddings and versioning, hybrid BM25 plus vector retrieval with RRF fusion, cross-encoder reranking, grounded generation with forced citations, abstention and a fallback ladder, ACL enforcement during retrieval, and the insert/edit/delete freshness lifecycle. It then gives a stage-by-stage evaluation plan (chunk answerability, hit-rate@K, recall@K, nDCG, faithfulness, citation quality, end-to-end success), how to bootstrap labels with no gold set, and follow-ups on debugging a wrongly cited document, embedding-model upgrades, multi-hop versus list-all queries, near-duplicates, conflicting sources and validating an LLM judge.

Read the full OpenAI Machine Learning Engineer interview experience this question came from

Community answers

Answer by Jack

chunking is also important to cover
|Home/ML System Design/OpenAI
OpenAI logo
OpenAI
Dec 15, 2025
hardMachine Learning EngineerOnsiteML System Design
78
0
Question

You are building a Retrieval-Augmented Generation (RAG) system for question answering over a private corpus — engineering wikis, design docs, runbooks, support tickets, PDFs, knowledge-base articles. Users ask natural-language questions in an interactive chat surface and expect grounded answers with citations back to the source documents.

Design the end-to-end system, describe the optimizations that make it high-quality, low-latency and trustworthy in production, and — for every component — explain how you would evaluate it. The evaluation half is not an afterthought here: for each stage the interviewer wants a named failure mode, the metric that detects it, and how you would get the labels.

Work through the following parts:

  1. End-to-end architecture and data flow — the full component set (ingestion, parsing, chunking, embeddings, indexing, retrieval, reranking, generation) and the split between the asynchronous indexing path and the synchronous query path.
  2. Ingestion and preprocessing — connectors, change detection, and layout-aware parsing of heterogeneous formats (PDF, HTML, wiki, Markdown), including tables and scanned documents.
  3. Chunking — splitting strategy for long, structured documents; size and overlap trade-offs; what metadata rides along with each chunk.
  4. Embeddings and indexing — embedding choice and versioning, ANN index, a parallel lexical (BM25) index, and metadata/ACL-filterable fields.
  5. Retrieval — query understanding and rewriting, top-N, metadata filters, hybrid sparse+dense search and fusion.
  6. Reranking — whether to add a second stage, what it buys, and what it costs.
  7. Context assembly and generation — prompt construction, dedupe and ordering, context compression, citation format.
  8. Grounding, guardrails and fallback — how you reduce hallucinations, and exactly what the system does when retrieval comes back weak or the question is ambiguous.
  9. Continuous ingestion and freshness — inserts, edits , and deletions , and the guarantee that deleted or superseded content can never be retrieved or cited.
  10. Access control — enforcing per-user ACLs without leaking even the existence of a forbidden document.
  11. Latency and cost — how you hit the interactive latency budget and where you trade quality for latency or cost.
  12. Evaluation plan, stage by stage — ingestion/chunking quality, retrieval quality, reranking quality, generation quality and grounding, and end-to-end user success. Cover both offline and online evaluation, and say how you obtain labels if no gold set exists.
  13. Monitoring, drift and the improvement loop — what you log, what you alert on, how you detect corpus/query drift, and how production failures become regression tests.

Constraints & assumptions

State your own numbers explicitly, but a reasonable baseline to design against:

  • Corpus on the order of 10^6–10^7 chunks after splitting, heterogeneous formats, long documents.
  • Interactive latency target: p95 end-to-end < 2 s to first token, with answer streaming.
  • Ingestion freshness SLA: a new, edited or deleted document is reflected in retrieval within minutes , not hours.
  • Documents carry access-control metadata (ACLs) — not every user may see every document.
  • Citations are mandatory ; abstention ("I couldn't find this in the docs") is allowed.
  • You may assume access to a managed or self-hosted vector index , a lexical/BM25 index , an embedding model, a cross-encoder reranker, and a generation LLM.

Clarifying questions to ask Guidance

A strong candidate scopes the problem before designing:

  • What is the expected query volume (QPS) and the read/write ratio (query rate vs. ingestion rate)?
  • What is the query mix — single-fact lookup, multi-hop reasoning, summarization, or "list all X"? This decides chunk size, top-k, and whether you need iterative retrieval.
  • Are answers single-turn or conversational (must the retriever resolve follow-up references against chat history)?
  • How strict are the access-control and data-isolation requirements — must we prevent even leaking the existence of a forbidden document?
  • What is the cost budget per query (embedding calls, reranker calls, generation tokens)?
  • How will quality be judged — is there an existing labeled eval set , or do we need to bootstrap one?
  • What is the output contract — free text or structured JSON, and what is the tolerance for abstention versus always producing something?

Follow-up questions Guidance

  • A user reports the system confidently cited the wrong document . Walk through how you localize the failure — chunk never retrieved, retrieved but mis-ranked, or present in the prompt and ignored — and what you change for each.
  • Your embedding model is upgraded to a new version. What is your re-indexing and rollout plan, and how do you avoid a mixed-embedding-space index where old and new vectors are incomparable?
  • How would you extend the system to handle multi-hop questions , and how does that differ from "list all / aggregate" questions?
  • The corpus contains near-duplicate documents (e.g. copies of the same runbook). How does this hurt retrieval, and how would you handle it?
  • Two retrieved documents contradict each other . What should the system do?
  • You use an LLM as a judge in your offline eval. How do you know the judge is any good, and what would make you stop trusting it?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...