Evaluate Instagram's Short-Video Recommender System Success

Quick Overview

Evaluates success metrics and experimentation for Instagram's short-video recommender feed. Strong answers choose a durable user-value metric, understand heavy-tailed engagement distributions, handle metric trade-offs, and outline a safe A/B test with retention, safety, creator, and app-level guardrails.

Evaluate Instagram's Short-Video Recommender System Success

Company: Meta

Role: Data Scientist

Category: Analytics & Experimentation

Difficulty: medium

Interview Round: Onsite

##### Scenario Instagram is launching a short-video recommender feed. ##### Question a) Which single metric would you use to evaluate the recommendation system’s success and why? b) Sketch or describe the expected distribution of that metric, labelling median, mode, and 95th percentile. c) Metric A rises while Metric B falls—what do you do? d) List the end-to-end A/B-testing steps for this launch. ##### Hints Discuss engagement vs retention, heavy-tailed distributions, guardrail metrics, power calculation and success criteria.

Quick Answer: Evaluates success metrics and experimentation for Instagram's short-video recommender feed. Strong answers choose a durable user-value metric, understand heavy-tailed engagement distributions, handle metric trade-offs, and outline a safe A/B test with retention, safety, creator, and app-level guardrails.

|Home/Analytics & Experimentation/Meta
Meta logo
Meta
Jul 12, 2025, 6:59 PM
mediumData ScientistOnsiteAnalytics & Experimentation
110
0

Evaluate Instagram's Short-Video Recommender System Success

Instagram is launching a short-video recommender feed. You are asked to choose metrics, reason about metric distributions and trade-offs, and outline the A/B test plan for launch.

Constraints & Assumptions

  • Treat this as an evaluation plan for a recommender/ranking system, not model architecture design.
  • Assume the product logs impressions, plays, watch time, skips, likes, shares, follows, hides, reports, sessions, and retention.
  • The system should improve user value while protecting content quality, safety, creator ecosystem health, and app-level engagement.
  • Short-video engagement metrics are likely heavy-tailed and should be analyzed carefully.

Clarifying Questions to Ask Guidance

  • Is the short-video feed a new surface or a replacement for an existing ranking model?
  • What is the product goal: retention, meaningful entertainment, creator discovery, revenue, or engagement?
  • Are there constraints around content safety, teenage users, or creator distribution?
  • What time horizon should determine launch: same-session, weekly, or long-term retention?

Part 1 - Choose a Primary Metric

Which single metric would you use to evaluate the recommendation system's success, and why?

What This Part Should Cover Guidance

  • A defensible primary metric such as retained short-video watch time per eligible active user, quality watch time, or weekly engaged users.
  • Why the metric reflects user value better than raw views alone.
  • Guardrails for hides, reports, skips, session abandonment, retention, creator concentration, and app cannibalization.

Part 2 - Describe the Distribution

Sketch or describe the expected distribution of the primary metric, labeling the median, mode, and 95th percentile.

What This Part Should Cover Guidance

  • Heavy-tailed user-level engagement, many low-activity users, and a small number of very high-usage users.
  • Difference between mean, median, mode, and percentile-based views of performance.
  • Whether to cap, winsorize, transform, or use robust analyses.

Part 3 - Resolve Metric Trade-offs

If Metric A rises while Metric B falls, what would you do?

What This Part Should Cover Guidance

  • Identify which metric is primary and which are guardrails.
  • Segment and cohort analysis to detect whether benefits are concentrated or harms are broad.
  • Examples such as watch time rising while retention, satisfaction, creator diversity, or reports worsen.
  • Decision rules for launch, rollback, or iteration.

Part 4 - Outline the A/B Test

List the end-to-end A/B testing steps for this launch.

What This Part Should Cover Guidance

  • Hypothesis, eligibility, randomization unit, sample size and power, ramp plan, duration, instrumentation validation, and analysis.
  • Primary metric, secondary metrics, guardrails, heterogeneity analysis, and novelty effects.
  • Launch criteria, rollback plan, and monitoring after launch.

What a Strong Answer Covers Guidance

A strong answer chooses a metric tied to durable user value, understands heavy-tailed recommender metrics, handles trade-offs explicitly, and describes a safe experiment and rollout plan.

Follow-up Questions Guidance

  • How would you detect that the recommender is creating unhealthy usage?
  • What if total watch time rises but app-level retention falls?
  • How would you measure creator ecosystem fairness?
Loading comments...