Assessing whether a new metric A is meaningful for News Feed
Company: Meta
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Technical Screen
Scenario: Another team proposes metric A as a better proxy for meaningful interactions. Evaluate its predictive power, actionability, and collection cost before adoption.
Question 1: How would you assess whether metric A is meaningful for News Feed? (Hint: predictive power, actionability, collection cost)
Quick Answer: Evaluates whether a proposed News Feed proxy metric is meaningful for long-term user value. Strong answers test predictive power, actionability, gaming risk, true-outcome alignment, collection cost, privacy, latency, and decision criteria.
A partner team proposes metric A as a proxy for "meaningful interactions" in News Feed. Before adopting it, evaluate whether it is a good proxy for underlying product goals such as user value, long-term engagement, and satisfaction.
Assess metric A across predictive power, actionability, and collection cost.
Constraints & Assumptions
Define the true target outcomes before judging the proxy.
Avoid adopting a metric that can be gamed or creates perverse incentives.
Include engineering, privacy, latency, and experimentation costs.
Provide decision criteria, not just a discussion.
Clarifying Questions to Ask Guidance
What exactly is metric A and how is it logged?
Is A intended for ranking, monitoring, experimentation, or diagnostics?
What true outcomes should A predict and at what time horizon?
What segments and surfaces are in scope?
Part 1 - Predictive Power
Does A reliably predict outcomes we care about?
What This Part Should Cover Guidance
Test correlation and incremental predictive value for retention, satisfaction, survey quality, meaningful engagement, safety, and long-term value.
Use out-of-sample validation, time splits, segment analysis, and calibration.
Check whether A leads outcomes rather than only reflecting them.
Part 2 - Actionability
Can product or ranking move A in the right direction, and will that improve true goals?
What This Part Should Cover Guidance
Run experiments or ranking simulations that optimize A and observe true outcomes.
Look for gaming, clickbait, low-quality engagement, or harmful side effects.
Include guardrails for safety, trust, creator health, fairness, and retention.
Part 3 - Collection Cost
What are the costs of logging and using A at scale?
What This Part Should Cover Guidance
Assess instrumentation, infrastructure, latency, storage, privacy, data quality, and experiment-readout cost.
Decide whether A should be a ranking feature, a diagnostic metric, or rejected.
Follow-up Questions Guidance
What if A predicts retention but worsens survey satisfaction?