Enforce Tenant and Session Isolation in an Agent RAG System
Company: Qualified Health
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
A multi-tenant Agent/RAG service stores many users' document vectors in shared infrastructure and answers questions using retrieved content and optional external tools. Design the authorization and guardrail boundaries that prevent cross-tenant or cross-session data exposure.
Explain how traditional gateway controls differ from retrieval authorization, model-level checks, and tool-execution policy. Uploaded text may contain instructions attempting to redirect the agent.
### Constraints & Assumptions
- A shared vector cluster does not imply shared access to every vector.
- The source calls for defenses against prompt injection, sensitive-data exposure, jailbreak attempts, and abusive request rates.
- Model-generated arguments and uploaded document contents are untrusted inputs to authorization decisions.
- This design does not assume any classifier perfectly detects malicious text.
### Clarifying Questions to Ask
- Is access scoped by tenant, user, session, individual document, or a combination?
- Which tools can change external state, and which require explicit user authorization?
- Can collaborators share a session, and how are revocation and document expiry propagated?
- What content may enter model prompts, provider logs, result caches, and application telemetry?
### Part 1 — Enforce Data Isolation
Trace identity and permissions through upload, indexing, retrieval, context assembly, and answer delivery. Explain how caching and asynchronous jobs preserve the same boundary.
#### What This Part Should Cover
- Trusted ownership metadata and server-derived authorization filters.
- Post-retrieval enforcement or defense in depth before context reaches a model.
- Scoped caches, object access, expiry, and permission-version handling.
### Part 2 — Place Guardrails at the Right Layer
Assign responsibilities to the gateway, orchestrator, retrieval service, model checks, and tool broker. Explain how a malicious instruction inside a retrieved document is prevented from authorizing an unrelated action.
#### What This Part Should Cover
- Authentication and resource limits versus semantic content checks.
- Treating retrieved text as data, not authority.
- Least-privilege tool execution with explicit intent checks, validation, and auditability.
```hint The model does not choose its tenant
A generated tool argument naming a different session or tenant must not replace the authenticated scope carried by the server.
```
### What a Strong Answer Covers
- End-to-end ownership enforcement before sensitive content or tool authority is exposed.
- Layered controls with different failure modes, rather than reliance on one prompt or classifier.
- Consistent protection across retrieval, caches, tools, logs, and permission changes.
### Follow-up Questions
- Why is filtering the answer after generation too late to prevent a cross-tenant context leak?
- How could a globally shared semantic-answer cache bypass otherwise correct vector filters?
- What happens when a document's access is revoked while an asynchronous indexing job is still running?
Overview: Isolate tenants and sessions across RAG retrieval, caches, prompts, and tools, with server-enforced authorization and layered prompt-injection defenses.
Read the full Qualified Health Software Engineer interview experience this question came from