Design an Enterprise Agent Platform with Guardrails and Latency Control
Company: Capital One
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design an enterprise agent platform. Explain the platform's execution architecture, how you choose between an agent and a predefined workflow, where guardrails operate, and how you control end-to-end latency.
### Constraints & Assumptions
- The reported technical topics are enterprise agent-platform design, guardrails, agent versus workflow, and latency. No original workload, deployment scale, tool catalog, or latency target is supplied.
- **Practice scope:** enterprise applications submit requests under an authenticated user and organization. A run may retrieve permitted enterprise information and invoke registered tools. Treat read-only lookup and an action that changes an enterprise record as contrasting examples, not as reported business requirements.
- **Practice security boundary:** retrieved content and model output cannot grant permissions. The platform must independently authorize tool actions. Discuss where a proposed write requires human approval rather than assuming all tools can act autonomously.
- State your workload and failure assumptions before choosing capacity, persistence, or deployment details. Do not invent a numerical service objective.
### Clarifying Questions to Ask
- Which steps require open-ended planning, and which have a known sequence and exact validation rules?
- Which tools read data, which change it, and whose permissions apply during a run?
- What must survive a worker crash, and which external actions support idempotency or status reconciliation?
- Is the latency goal time to first useful response, final completion, or both, and how much of it may be spent waiting for approval?
### Part 1 — Define the Platform and Execution State
Describe request admission, run state, orchestration, model access, tool execution, and operational visibility. Trace a request that retrieves information and then proposes an enterprise-record update.
#### What This Part Should Cover
- Identity and authorization context that remains attached to each run and tool call.
- Durable run/step state and the boundary around external effects.
- Versioned tool contracts and enough trace information to diagnose a failed or slow run.
### Part 2 — Choose Agents and Workflows Deliberately
Explain how a predefined workflow differs from a model-directed agent. Show where each fits in the proposed platform and how a hybrid execution limits uncontrolled planning.
#### What This Part Should Cover
- Who chooses the next step and what state transitions are allowed.
- Reproducibility and validation advantages of fixed workflows versus the flexibility and uncertainty of model-directed plans.
- Explicit stopping conditions for repeated tool calls or unproductive replanning.
### Part 3 — Put Guardrails at Enforceable Boundaries
Explain how the platform handles retrieved instructions that conflict with the user's request, a proposed tool call outside the user's permissions, and a write waiting for human approval.
#### What This Part Should Cover
- Separation of untrusted data from authority and independent authorization at tool execution.
- Validation before effects, approval bound to a specific proposed action, and checks before output reaches the user.
- A distinction between probabilistic model checks and deterministic access control.
### Part 4 — Control Latency Without Weakening the Contract
Break down the latency of a run and propose changes that reduce avoidable delay. Explain which steps can run in parallel and which must wait for dependencies or approval.
#### What This Part Should Cover
- Measurement across admission, model calls, retrieval, tools, and delivery.
- Safe caching, constrained parallelism, and model/routing choices evaluated against task quality.
- Deadline and cancellation behavior, including external actions whose outcomes are uncertain.
```hint Locate the authority boundary
A model can propose an action and retrieved content can describe one. Decide which component has the authority to permit it and what exact arguments an approval covers.
```
### What a Strong Answer Covers
- A concrete run lifecycle connecting orchestration, tool contracts, identity, and durable state.
- A justified choice of workflow, agent, or hybrid rather than treating every task as open-ended planning.
- Enforceable guardrails and a latency plan that preserves permission and approval checks.
### Follow-up Questions
- What happens if a write succeeds remotely but the worker crashes before recording its result?
- How would a changed tool schema or revoked permission affect a paused run awaiting approval?
- When does parallel retrieval reduce latency, and when can extra model or tool calls make both latency and cost worse?
Overview: Design an enterprise agent platform with durable execution, deliberate workflow-versus-agent choices, enforceable tool guardrails, and a dependency-aware latency plan.
Read the full Capital One Software Engineer interview experience this question came from