Design a RAG system end to end

Quick Overview

This question evaluates a candidate's ability to design a production-grade Retrieval‑Augmented Generation (RAG) system, testing competencies in scalable ML system architecture, embedding and vector retrieval strategies, prompt orchestration, freshness and latency engineering, security/access controls, and evaluation metrics.

Design a RAG system end to end

Company: Amazon

Role: Applied Scientist

Category: ML System Design

Difficulty: hard

Interview Round: Technical Screen

Design a retrieval-augmented generation system for enterprise text. Specify the ingestion pipeline (chunking, embedding generation, indexing), retrieval strategy (vector search, hybrid retrieval, reranking), prompt orchestration, grounding and citations, freshness handling, latency/throughput targets, and privacy controls. Discuss evaluation for relevance and answer quality, approaches to reduce hallucinations, and how you would scale and monitor the system in production.

Overview: This question evaluates a candidate's ability to design a production-grade Retrieval‑Augmented Generation (RAG) system, testing competencies in scalable ML system architecture, embedding and vector retrieval strategies, prompt orchestration, freshness and latency engineering, security/access controls, and evaluation metrics.

Community answers

Answer by Lou

A production enterprise RAG system should treat retrieval, authorization, freshness, and grounding as first-class services—not simply “vector database + LLM.” The central invariant is: a user may retrieve and cite only content they are authorized to see, and the model must abstain when the retrieved evidence is insufficient. Use independently scalable services: • Source connectors: Wiki, document management system, ticketing, email archive, file shares, and approved enterprise SaaS APIs. • Document-processing pipeline: File-type routing, text extraction, OCR, table-aware parsing, language detection, PII/sensitivity classification, deduplication, chunking, and metadata/ACL propagation. • Indexing plane: A lexical index plus an approximate-nearest-neighbor vector index, both keyed by immutable content version and tenant. • Online query plane: Authentication, authorization filtering, retrieval, reranking, prompt construction, LLM inference, citation validation, streaming response. • Control and observability plane: Evaluation, audit, lineage, red-team testing, feature flags, index health, quality monitoring, cost and latency dashboards. For multi-tenancy, use logical tenant isolation at minimum and physical/index isolation for high-sensitivity or regulated tenants. Every document and chunk carries a  tenant_id ,  document_id ,  version_id , security labels, ACL principal/group references, source URI, timestamps, checksum, parser provenance, and deletion state.
|Home/ML System Design/Amazon
Amazon logo
Amazon
Sep 6, 2025
hardApplied ScientistTechnical ScreenML System Design
26
0

Design a Retrieval‑Augmented Generation (RAG) System for Enterprise Text

Context

You are building a production RAG system that answers employee questions using internal enterprise text (wikis, PDFs, tickets, emails, docs). Data is sensitive and access-controlled. Assume multi-tenant use, mixed document formats, English-first, with the following baseline constraints:

  • Corpus: 5–10 million pages, tens of millions of chunks.
  • Traffic: 200 QPS peak; target end-to-end p95 latency ≤ 2.0 s with server-streamed tokens.
  • Freshness: new or updated content should be searchable within 15 minutes.

Tasks

Design the system and specify:

  1. Ingestion pipeline: chunking strategy, embedding generation, and indexing.
  2. Retrieval strategy: vector search, hybrid retrieval, and reranking.
  3. Prompt orchestration: how the LLM is instructed and grounded; how citations are produced.
  4. Freshness handling: incremental updates, cache invalidation, time-aware ranking.
  5. Latency and throughput targets with a rough budget.
  6. Privacy and security controls for enterprise data.
  7. Evaluation: measuring relevance and answer quality; datasets and metrics.
  8. Reducing hallucinations: techniques across retrieval and generation.
  9. Scale and monitoring: how you would scale, operate, and observe the system in production.

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...