System design round
They have you design an eval system to evaluate the results of an Agent using different models, e.g. tool usage, response speed, plus testing different system prompts, that kind of thing.
The original question is below:
Design a system for evaluating AI agents. Users of the system are going to be teams developing AI agents, e.g. agents which automatically analyze contact center conversations to produce a report. They need a platform which allows evaluating how well their agents are performing. E.g. they need to make a change to the system prompt of an AI agent, or the model it uses, and measure how this change impacts the agent's performance.
Discussion
Loading comments…