Define Success Metrics for Euro-Chat Customer-Service Chatbot
Company: Meta
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Onsite
##### Scenario
An e-commerce company deploys euro-chat chatbot to handle B2C customer-service inquiries.
##### Question
How would you define a primary success metric for the euro-chat customer-service chatbot? What guardrail metrics would you track? Describe an experiment you would run to test whether the chatbot improves customer experience.
##### Hints
Consider resolution rate, CSAT and handle time. Outline control/treatment, randomisation unit, success threshold, sample sizing, and monitoring.
Quick Answer: Meta customer-service chatbot analytics prompt covering primary success metrics, verified resolution, CSAT, recontact, guardrails, A/B test design, safety, escalation, and monitoring.
Success Metrics for the Euro-Chat Customer-Service Chatbot
An e-commerce company deploys a customer-service chatbot called euro-chat to handle B2C support inquiries across web or app chat. The bot can answer some questions and escalate to human agents when needed.
Constraints & Assumptions
Define one primary success metric that captures customer outcome, not just bot containment.
Include guardrails for experience quality, safety, and operations.
Design an experiment to test whether the chatbot improves customer experience.
Some intents may be too risky for bot-only handling and should be excluded or routed to humans.
Clarifying Questions to Ask Guidance
Which intents are eligible for the bot at launch?
What is the current human-agent baseline for resolution, CSAT, handle time, and recontact?
Can customers self-report whether the issue was solved?
Are we randomizing by user, conversation, intent, retailer, or market?
What a Strong Answer Covers Guidance
A primary metric such as verified resolution without recontact within 7 days, or customer-verified resolution rate.
Clear denominator, measurement window, intent eligibility, and handling of missing survey responses.
Guardrails: CSAT, recontact rate, escalation rate, time to resolution, abandonment, hallucination/policy-error rate, PII/safety incidents, refunds, support cost, and latency.
Experiment design: control versus treatment, randomization unit, stratification by intent, sample size/MDE, monitoring, stop criteria, and ramp.
Analysis plan using intent-to-treat, confidence intervals, segment cuts, and separate analysis for high-risk intents.
Practical caveats around nonresponse bias, bot fallback quality, agent capacity, and customer self-selection.
Follow-up Questions Guidance
Why is containment rate alone a bad primary metric?
How would you handle urgent or high-risk support cases?