Answer Enterprise Questions Across Internal and External Sources
Company: Cohere
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design an AI system that helps enterprise users answer complex questions by searching internal and external information sources, including business systems such as Salesforce.
Explain how questions are decomposed into retrieval work, how evidence from different sources is combined, and how the answer preserves user permissions, provenance, and uncertainty.
### Constraints & Assumptions
- The source supplies enterprise question answering across internal and external sources; it does not define scale, model, latency, connector APIs, or a specific company architecture.
- **Practice scope:** the system reads authorized documents or records and returns an evidence-grounded answer. Writing to business systems is not required by the source.
- Different sources can have different schemas, freshness, permissions, rate limits, and availability.
- A source name alone does not imply that every user may access all records in it.
### Clarifying Questions to Ask
- Which questions require structured record queries, unstructured search, or both?
- Must retrieval reflect live source permissions and updates, or is a bounded-staleness index acceptable?
- Should the system ask for clarification when an entity or business term is ambiguous?
- What evidence and source links should accompany the answer, and how should unavailable sources be reported?
### Part 1 — Connect and Retrieve
Design connectors, indexing or live query paths, identity propagation, and source selection.
#### What This Part Should Cover
- Source-specific schemas and permissions carried into retrieval.
- An explicit trade-off between fresh live reads and indexed search.
- Bounded query plans and safe interpretation of generated query arguments.
### Part 2 — Synthesize a Complex Answer
Describe evidence normalization, multi-step retrieval, entity resolution, and answer generation.
#### What This Part Should Cover
- Stable provenance and time/version context for each result.
- No unsupported joining of similarly named entities or incompatible measures.
- Handling missing, contradictory, or insufficient evidence without inventing facts.
### Part 3 — Operate Reliably
Explain source failures, quota pressure, caching, and evaluation of answer quality.
#### What This Part Should Cover
- Bounded retries and clear partial-answer semantics.
- Authorization- and freshness-aware caches.
- Evaluation that distinguishes retrieval coverage from supported synthesis.
```hint Two matching names need not identify one account
Before combining records from different systems, decide which identifiers or mapping evidence establish that they refer to the same entity.
```
### What a Strong Answer Covers
- A complete cross-source read-and-answer architecture.
- Permissions enforced before evidence is exposed to the model.
- Source-bound synthesis with explicit uncertainty and operational limits.
### Follow-up Questions
- How would a permission revocation affect an already indexed record and a cached answer?
- What should happen if one source's data is current and another source's snapshot is several days old?
- How would you detect an answer that cites relevant sources but makes a claim none of them supports?
Overview: Design enterprise AI question answering across internal and external sources with permission-aware retrieval, entity mapping, evidence-grounded synthesis, and freshness controls.