Evaluate Factual Accuracy and Consistency in a Domain-Specific RAG Chatbot

Read the full interview experience this question came from →

Quick Overview

Evaluate a rare-disease RAG chatbot using domain consistency and expert review, and distinguish factual accuracy from retrieval metrics and lexical overlap.

Evaluate Factual Accuracy and Consistency in a Domain-Specific RAG Chatbot

Company: C3 AI

Role: Software Engineer

Category: Machine Learning

Difficulty: medium

Interview Round: Online Assessment

A retrieval-augmented generation (RAG) chatbot assists doctors in investigating rare diseases. Factual accuracy and consistency are important, but retrieval metrics such as signal-to-noise ratio and recall may not adequately evaluate the answers the system generates. Select the **two** approaches below that most directly address this evaluation gap. Explain why they complement each other and why the other options are weaker answers to this particular evaluation question. 1. Develop a domain-specific metric that assesses the coherence and plausibility of proposed diagnoses against a medical knowledge graph. 2. Combine traditional RAG metrics with human review by medical experts to assess factual accuracy and potential biases. 3. Use active learning so feedback on generated diagnoses continually modifies the pipeline, with the aim of improving specificity and reducing false positives. 4. Use similarity metrics such as ROUGE to assess the relevance of the generated output. ### Constraints & Assumptions - This is a question about evaluating an information system, not making a patient diagnosis or choosing treatment. - The system retrieves evidence and generates an answer; assess those stages separately as well as together. - Selecting an evaluation method does not establish that the system is clinically reliable or that a knowledge graph is complete. - The source supplies no dataset, metric weights, performance target, or measured results. Do not invent them. ### Clarifying Questions to Ask - Does the evaluation set contain expert-checked evidence and accepted alternative answers, or only one reference wording per case? - Which kinds of inconsistency should be measured: contradiction with retrieved evidence, conflict with domain knowledge, or variation across repeated equivalent queries? - How will reviewer disagreement and gaps in knowledge-graph coverage be recorded rather than hidden in a single score? ```hint Distinguish measurement from modification One option changes the pipeline in response to feedback. Decide whether that alone provides an independent way to measure the current system's factual correctness. ``` ### What a Strong Answer Covers - Selection of two complementary approaches that evaluate domain consistency and factual accuracy beyond retrieval quality. - The limits of knowledge-graph plausibility and the role of expert adjudication. - Why active learning and lexical overlap do not independently establish the desired answer quality. - Separation of retrieval failure, unsupported generation, missing evidence, and unresolved expert disagreement. ### Follow-up Questions - How could a response achieve high retrieval recall yet still contain an unsupported claim? - Why might an accurate paraphrase score poorly on ROUGE while an incorrect answer shares many words with a reference? - How would you keep feedback used to improve the model separate from the evidence used to evaluate that improvement?

Overview: Evaluate a rare-disease RAG chatbot using domain consistency and expert review, and distinguish factual accuracy from retrieval metrics and lexical overlap.

Read the full C3 AI Software Engineer interview experience this question came from

|Home/Machine Learning/C3 AI
C3 AI logo
C3 AI
Jan 23, 2026
mediumSoftware EngineerOnline AssessmentMachine Learning
0
0

A retrieval-augmented generation (RAG) chatbot assists doctors in investigating rare diseases. Factual accuracy and consistency are important, but retrieval metrics such as signal-to-noise ratio and recall may not adequately evaluate the answers the system generates.

Select the two approaches below that most directly address this evaluation gap. Explain why they complement each other and why the other options are weaker answers to this particular evaluation question.

  1. Develop a domain-specific metric that assesses the coherence and plausibility of proposed diagnoses against a medical knowledge graph.
  2. Combine traditional RAG metrics with human review by medical experts to assess factual accuracy and potential biases.
  3. Use active learning so feedback on generated diagnoses continually modifies the pipeline, with the aim of improving specificity and reducing false positives.
  4. Use similarity metrics such as ROUGE to assess the relevance of the generated output.

Constraints & Assumptions

  • This is a question about evaluating an information system, not making a patient diagnosis or choosing treatment.
  • The system retrieves evidence and generates an answer; assess those stages separately as well as together.
  • Selecting an evaluation method does not establish that the system is clinically reliable or that a knowledge graph is complete.
  • The source supplies no dataset, metric weights, performance target, or measured results. Do not invent them.

Clarifying Questions to Ask Guidance

  • Does the evaluation set contain expert-checked evidence and accepted alternative answers, or only one reference wording per case?
  • Which kinds of inconsistency should be measured: contradiction with retrieved evidence, conflict with domain knowledge, or variation across repeated equivalent queries?
  • How will reviewer disagreement and gaps in knowledge-graph coverage be recorded rather than hidden in a single score?

What a Strong Answer Covers Guidance

  • Selection of two complementary approaches that evaluate domain consistency and factual accuracy beyond retrieval quality.
  • The limits of knowledge-graph plausibility and the role of expert adjudication.
  • Why active learning and lexical overlap do not independently establish the desired answer quality.
  • Separation of retrieval failure, unsupported generation, missing evidence, and unresolved expert disagreement.

Follow-up Questions Guidance

  • How could a response achieve high retrieval recall yet still contain an unsupported claim?
  • Why might an accurate paraphrase score poorly on ROUGE while an incorrect answer shares many words with a reference?
  • How would you keep feedback used to improve the model separate from the evidence used to evaluate that improvement?
Loading comments...