Design an Evaluation Platform for AI Agents Across Prompt and Model Changes

Read the full interview experience this question came from →

Quick Overview

An ML system design question about a platform that lets teams evaluate AI agents and measure how a change of system prompt or model affects quality, tool use and response speed. It tests evaluation design, scoring with references and LLM judges, statistical comparison of versions and scalable execution.

Design an Evaluation Platform for AI Agents Across Prompt and Model Changes

Company: Cresta

Role: Software Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

Design a system for evaluating AI agents. The users of the system are teams that develop AI agents, for example agents that automatically analyze contact center conversations to produce a report. These teams need a platform that lets them evaluate how well their agents are performing. For example, when a team changes an agent's system prompt, or the model the agent uses, it needs to measure how that change affects the agent's performance. The platform should let teams compare agent configurations along several dimensions, such as how well the agent uses its tools, how quickly it responds, and how different system prompts and different models change the results. ```hint Start from the unit of comparison Pin down what one evaluation run compares, such as a baseline agent configuration against a candidate on the same inputs, before designing storage or scoring. ``` ```hint Scoring without one right answer A report written from a conversation rarely has a single correct output. Decide which checks can be deterministic, which need a reference, which need a judge, and how you would come to trust that judge. ``` ### Clarifying Questions - What does an agent look like to the platform: a configuration (model, system prompt, tools) that the platform runs itself, or an endpoint that each team hosts? - Is the focus offline evaluation on curated datasets, online evaluation on production traffic, or both? - Where do test inputs come from: curated conversations, samples of production traffic, synthetic data? Are there privacy restrictions on real conversations? - Is there ground truth, such as human-written reports or labeled fields, or must quality be judged without references? - Besides quality, which metrics matter: tool-use correctness, latency, cost? - Should an evaluation run be able to gate a prompt or model change before it ships, and how quickly must it finish? ### What a Strong Answer Covers - Core entities and versioning: agent versions (model, parameters, system prompt, tools), datasets and test cases, evaluators, runs and per-case results, with reproducibility. - Execution at scale: running agent versions over datasets with rate limits, retries, caching, controlled tool behavior, and repeated runs to account for nondeterminism. - Scoring: deterministic checks, reference-based metrics, LLM judges with rubrics calibrated against human labels, tool-trajectory evaluation, and latency and cost from traces. - Statistically sound comparison of a candidate against a baseline, with slicing and regression detection. - The product surface: APIs, reports with case-level drill-down, CI gating, and a human review loop. - Online evaluation, data privacy and multi-tenancy, and cost control. ### Follow-up Questions - How would you validate an LLM judge, and detect that its behavior changed after the judge model was upgraded? - How would you evaluate multi-turn agents whose later tool calls depend on the results of earlier ones? - How would you keep evaluation cost under control as datasets and the number of teams grow? - How would production failures flow back into the evaluation datasets?

Overview: An ML system design question about a platform that lets teams evaluate AI agents and measure how a change of system prompt or model affects quality, tool use and response speed. It tests evaluation design, scoring with references and LLM judges, statistical comparison of versions and scalable execution.

Read the full Cresta Software Engineer interview experience this question came from

|Home/ML System Design/Cresta
Cresta logo
Cresta
Sep 20, 2026
mediumSoftware EngineerOnsiteML System Design
0
0

Design a system for evaluating AI agents. The users of the system are teams that develop AI agents, for example agents that automatically analyze contact center conversations to produce a report. These teams need a platform that lets them evaluate how well their agents are performing. For example, when a team changes an agent's system prompt, or the model the agent uses, it needs to measure how that change affects the agent's performance.

The platform should let teams compare agent configurations along several dimensions, such as how well the agent uses its tools, how quickly it responds, and how different system prompts and different models change the results.

Clarifying Questions Guidance

  • What does an agent look like to the platform: a configuration (model, system prompt, tools) that the platform runs itself, or an endpoint that each team hosts?
  • Is the focus offline evaluation on curated datasets, online evaluation on production traffic, or both?
  • Where do test inputs come from: curated conversations, samples of production traffic, synthetic data? Are there privacy restrictions on real conversations?
  • Is there ground truth, such as human-written reports or labeled fields, or must quality be judged without references?
  • Besides quality, which metrics matter: tool-use correctness, latency, cost?
  • Should an evaluation run be able to gate a prompt or model change before it ships, and how quickly must it finish?

What a Strong Answer Covers Guidance

  • Core entities and versioning: agent versions (model, parameters, system prompt, tools), datasets and test cases, evaluators, runs and per-case results, with reproducibility.
  • Execution at scale: running agent versions over datasets with rate limits, retries, caching, controlled tool behavior, and repeated runs to account for nondeterminism.
  • Scoring: deterministic checks, reference-based metrics, LLM judges with rubrics calibrated against human labels, tool-trajectory evaluation, and latency and cost from traces.
  • Statistically sound comparison of a candidate against a baseline, with slicing and regression detection.
  • The product surface: APIs, reports with case-level drill-down, CI gating, and a human review loop.
  • Online evaluation, data privacy and multi-tenancy, and cost control.

Follow-up Questions Guidance

  • How would you validate an LLM judge, and detect that its behavior changed after the judge model was upgraded?
  • How would you evaluate multi-turn agents whose later tool calls depend on the results of earlier ones?
  • How would you keep evaluation cost under control as datasets and the number of teams grow?
  • How would production failures flow back into the evaluation datasets?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...