Design Metrics for Content Moderation and Chatbot Evaluation
Company: TikTok
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Technical Screen
##### Scenario
Trust & Safety data science: design metrics for a content-moderation A/B test and for evaluating a customer-service chatbot’s knowledge base.
##### Question
In a content-moderation A/B test where harmful-content prevalence is low, which short-term, user-centric metrics would you track to detect impact quickly and why? How would you design an experiment and select evaluation metrics to measure the quality and usefulness of a customer-service chatbot’s knowledge base?
##### Hints
Consider immediate user actions: report rates, dismissals, session exits, latency; for chatbot, precision/recall of answers, deflection rate, CSAT; discuss experiment design and trade-offs.
Quick Answer: This interview question evaluates metric design, causal reasoning, experiment setup, diagnostics, SQL/statistical checks, and recommendations in a realistic interview setting. A strong answer for Design Metrics for Content Moderation and Chatbot Evaluation states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.