Design Customer Support with Virtual-Agent-to-Human Escalation
Company: Google
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
Design a customer-support system in which a user converses with a virtual agent and can be escalated to a human support agent. Explain how conversation context, ownership, and message delivery are preserved during the handoff.
### Constraints & Assumptions
- The source supplies the virtual-agent-to-human support scenario, without scale, channel, staffing, or SLA details.
- **Practice scope:** a text conversation has durable messages, one current response owner, an escalation queue, and a human-agent interface. Escalation may be requested by the user or triggered by an explicit routing policy.
- Clarify any bot tools and permissions separately; the handoff must not silently repeat an external action already taken by the virtual agent.
- Human availability is limited. The system needs a visible state for waiting rather than pretending every escalation is immediately staffed.
### Clarifying Questions to Ask
- What triggers escalation, and can the user always request it directly?
- Which skills, language, priority, or availability rules choose a human agent?
- Can a user continue sending messages while waiting, and what may the bot do during that period?
- Which context may be shared with the human agent, and how are generated summaries distinguished from the original transcript?
### Part 1 — Conversation and Ownership Model
Describe message persistence, virtual-agent responses, escalation requests, queueing, human claim, and resolution.
#### What This Part Should Cover
- A durable conversation state machine with one current response owner.
- An explicit waiting state and idempotent escalation request.
- Message identity and ordering across reconnects and repeated sends.
### Part 2 — Transfer Context and Stop Conflicting Replies
Explain the handoff boundary, transcript and summary delivery, messages arriving during transfer, and a delayed bot response after human takeover.
#### What This Part Should Cover
- A versioned ownership transition and fencing of stale bot or agent writes.
- Original messages and tool outcomes available alongside any generated summary.
- No lost messages between the transferred snapshot and live updates.
### Part 3 — Recover and Operate
Describe human disconnects, reassignment, cancellation, and operational measurements.
#### What This Part Should Cover
- Durable queue and assignment state with controlled recovery.
- No silent conversation abandonment or duplicate ownership.
- Metrics for waiting, handoff correctness, and resolution without invented performance claims.
```hint Handoff is a state transition, not just a notification
Sending a transcript to a human does not stop an in-flight bot response or prevent another human from claiming the same conversation.
```
### What a Strong Answer Covers
- A complete virtual-to-human lifecycle with durable history and clear ownership.
- Consistent context transfer and fencing of stale responses.
- Visible waiting and recovery behavior when no human is immediately available.
### Follow-up Questions
- What should happen if the bot completes a tool action after escalation was requested?
- How can a human inspect what the bot actually did rather than relying only on its summary?
- How would a reassignment avoid losing messages sent while the previous human was disconnected?
Overview: Design virtual-agent support with reliable human escalation, durable conversation ownership, complete context transfer, and recovery from interrupted handoffs.
Read the full Google Software Engineer interview experience this question came from