Answer Enterprise Questions Across Internal and External Sources

Quick Overview

Design enterprise AI question answering across internal and external sources with permission-aware retrieval, entity mapping, evidence-grounded synthesis, and freshness controls.

Answer Enterprise Questions Across Internal and External Sources

Company: Cohere

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design an AI system that helps enterprise users answer complex questions by searching internal and external information sources, including business systems such as Salesforce. Explain how questions are decomposed into retrieval work, how evidence from different sources is combined, and how the answer preserves user permissions, provenance, and uncertainty. ### Constraints & Assumptions - The source supplies enterprise question answering across internal and external sources; it does not define scale, model, latency, connector APIs, or a specific company architecture. - **Practice scope:** the system reads authorized documents or records and returns an evidence-grounded answer. Writing to business systems is not required by the source. - Different sources can have different schemas, freshness, permissions, rate limits, and availability. - A source name alone does not imply that every user may access all records in it. ### Clarifying Questions to Ask - Which questions require structured record queries, unstructured search, or both? - Must retrieval reflect live source permissions and updates, or is a bounded-staleness index acceptable? - Should the system ask for clarification when an entity or business term is ambiguous? - What evidence and source links should accompany the answer, and how should unavailable sources be reported? ### Part 1 — Connect and Retrieve Design connectors, indexing or live query paths, identity propagation, and source selection. #### What This Part Should Cover - Source-specific schemas and permissions carried into retrieval. - An explicit trade-off between fresh live reads and indexed search. - Bounded query plans and safe interpretation of generated query arguments. ### Part 2 — Synthesize a Complex Answer Describe evidence normalization, multi-step retrieval, entity resolution, and answer generation. #### What This Part Should Cover - Stable provenance and time/version context for each result. - No unsupported joining of similarly named entities or incompatible measures. - Handling missing, contradictory, or insufficient evidence without inventing facts. ### Part 3 — Operate Reliably Explain source failures, quota pressure, caching, and evaluation of answer quality. #### What This Part Should Cover - Bounded retries and clear partial-answer semantics. - Authorization- and freshness-aware caches. - Evaluation that distinguishes retrieval coverage from supported synthesis. ```hint Two matching names need not identify one account Before combining records from different systems, decide which identifiers or mapping evidence establish that they refer to the same entity. ``` ### What a Strong Answer Covers - A complete cross-source read-and-answer architecture. - Permissions enforced before evidence is exposed to the model. - Source-bound synthesis with explicit uncertainty and operational limits. ### Follow-up Questions - How would a permission revocation affect an already indexed record and a cached answer? - What should happen if one source's data is current and another source's snapshot is several days old? - How would you detect an answer that cites relevant sources but makes a claim none of them supports?

Overview: Design enterprise AI question answering across internal and external sources with permission-aware retrieval, entity mapping, evidence-grounded synthesis, and freshness controls.

|Home/System Design/Cohere
Cohere logo
Cohere
Oct 1, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Design an AI system that helps enterprise users answer complex questions by searching internal and external information sources, including business systems such as Salesforce.

Explain how questions are decomposed into retrieval work, how evidence from different sources is combined, and how the answer preserves user permissions, provenance, and uncertainty.

Constraints & Assumptions

  • The source supplies enterprise question answering across internal and external sources; it does not define scale, model, latency, connector APIs, or a specific company architecture.
  • Practice scope: the system reads authorized documents or records and returns an evidence-grounded answer. Writing to business systems is not required by the source.
  • Different sources can have different schemas, freshness, permissions, rate limits, and availability.
  • A source name alone does not imply that every user may access all records in it.

Clarifying Questions to Ask Guidance

  • Which questions require structured record queries, unstructured search, or both?
  • Must retrieval reflect live source permissions and updates, or is a bounded-staleness index acceptable?
  • Should the system ask for clarification when an entity or business term is ambiguous?
  • What evidence and source links should accompany the answer, and how should unavailable sources be reported?

Part 1 — Connect and Retrieve

Design connectors, indexing or live query paths, identity propagation, and source selection.

What This Part Should Cover Guidance

  • Source-specific schemas and permissions carried into retrieval.
  • An explicit trade-off between fresh live reads and indexed search.
  • Bounded query plans and safe interpretation of generated query arguments.

Part 2 — Synthesize a Complex Answer

Describe evidence normalization, multi-step retrieval, entity resolution, and answer generation.

What This Part Should Cover Guidance

  • Stable provenance and time/version context for each result.
  • No unsupported joining of similarly named entities or incompatible measures.
  • Handling missing, contradictory, or insufficient evidence without inventing facts.

Part 3 — Operate Reliably

Explain source failures, quota pressure, caching, and evaluation of answer quality.

What This Part Should Cover Guidance

  • Bounded retries and clear partial-answer semantics.
  • Authorization- and freshness-aware caches.
  • Evaluation that distinguishes retrieval coverage from supported synthesis.

What a Strong Answer Covers Guidance

  • A complete cross-source read-and-answer architecture.
  • Permissions enforced before evidence is exposed to the model.
  • Source-bound synthesis with explicit uncertainty and operational limits.

Follow-up Questions Guidance

  • How would a permission revocation affect an already indexed record and a cached answer?
  • What should happen if one source's data is current and another source's snapshot is several days old?
  • How would you detect an answer that cites relevant sources but makes a claim none of them supports?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...