Evaluate Auto-Play Impact with Key Metrics and Experiment Design
Company: Uber
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Technical Screen
##### Scenario
Streaming platform considering auto-playing the next episode to improve engagement.
##### Question
What primary metrics and guardrail metrics would you track to evaluate auto-play? Design an experiment to measure the feature's impact; outline unit of randomization and duration. If churn increases in power users but total watch time rises overall, how would you decide whether to launch?
##### Hints
Discuss trade-offs, segmentation, north-star metric.
Evaluate Auto-Play Impact With Metrics and Experiment Design
A streaming platform is considering auto-playing the next episode to improve engagement. The company wants to know whether auto-play increases meaningful engagement without harming retention, user satisfaction, or platform health.
Constraints & Assumptions
Auto-play may increase watch time mechanically, so define meaningful engagement carefully.
Include primary metrics, guardrails, experiment design, and segment-level interpretation.
Consider user control, opt-out, and long-term retention.
Address the case where churn increases among power users but total watch time rises overall.
Clarifying Questions to Ask Guidance
What content types are eligible for auto-play?
Does auto-play apply to all profiles, including kids profiles?
Is the business optimizing retention, ad revenue, subscription value, or viewing satisfaction?
Can users disable auto-play?
What a Strong Answer Covers Guidance
Primary metric such as retention-adjusted engaged watch time, completed meaningful watches, or long-term LTV proxy.
Secondary metrics such as next-episode start rate, session length, completion rate, repeat visits, and content discovery.
Guardrails: churn, retention, opt-out/disable rate, skips within the first minute, complaints, ratings, rebuffering, latency, data/battery usage, content diversity, and ad fatigue.
Experiment design with user/account-level randomization, eligibility, exposure logging, stratification, sample size, duration, and triggered analysis.
Interpretation of heterogeneous effects: power-user churn may outweigh aggregate watch-time lift if it harms high-value retention or long-term health.
Decision framework using pre-registered thresholds, segment guardrails, and possible targeted rollout or opt-in design.
Follow-up Questions Guidance
Why is total watch time alone risky as a north-star metric?