Answer by Lou
A production enterprise RAG system should treat retrieval, authorization, freshness, and grounding as first-class services—not simply “vector database + LLM.” The central invariant is: a user may retrieve and cite only content they are authorized to see, and the model must abstain when the retrieved evidence is insufficient.
Use independently scalable services:
• Source connectors: Wiki, document management system, ticketing, email archive, file shares, and approved enterprise SaaS APIs.
• Document-processing pipeline: File-type routing, text extraction, OCR, table-aware parsing, language detection, PII/sensitivity classification, deduplication, chunking, and metadata/ACL propagation.
• Indexing plane: A lexical index plus an approximate-nearest-neighbor vector index, both keyed by immutable content version and tenant.
• Online query plane: Authentication, authorization filtering, retrieval, reranking, prompt construction, LLM inference, citation validation, streaming response.
• Control and observability plane: Evaluation, audit, lineage, red-team testing, feature flags, index health, quality monitoring, cost and latency dashboards.
For multi-tenancy, use logical tenant isolation at minimum and physical/index isolation for high-sensitivity or regulated tenants. Every document and chunk carries a tenant_id , document_id , version_id , security labels, ACL principal/group references, source URI, timestamps, checksum, parser provenance, and deletion state.