How would you evaluate stolen-post detection?
Company: Meta
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: hard
Interview Round: Technical Screen
You are interviewing for a Meta DSA (product analytics / data science) role. The product team is launching a new **Stolen Post Detection** algorithm that flags posts suspected of being copied/reposted without attribution, and then triggers actions (e.g., downrank, warning label, creator notification, or removal).
Design an evaluation plan covering:
1) **Problem diagnosis & clarification:** What questions would you ask to clarify the product goal and the meaning of “stolen” (e.g., exact duplicate vs paraphrase vs meme templates), enforcement actions, and success criteria?
2) **Harms & tradeoffs:** Enumerate likely failure modes and harms of false positives vs false negatives, including different stakeholder impacts (original creator, reposter, viewers, moderators).
3) **Metrics:** Propose a metric framework with (a) primary success metrics, (b) guardrails, and (c) offline model metrics. Include at least one metric that can move in opposite directions depending on threshold choice.
4) **Experiment design:** Propose an online experiment (or quasi-experiment if A/B is hard). Address logging, unit of randomization, interference/network effects, ramp strategy, and how you would compute/think about power/MDE.
5) **Post-launch monitoring:** What would you monitor to detect regressions or gaming, and how would you iterate on thresholds/policy over time?
Overview: This question evaluates product analytics, experimental design, and causal thinking for content-moderation algorithms—specifically metric specification, trade-off/harm analysis, and online experiment logistics—and is commonly asked to gauge a data scientist’s ability to balance detection accuracy, stakeholder impacts, and business objectives in production features; it is in the Analytics & Experimentation category for a Data Scientist position. At a high abstraction level it probes system-level reasoning around problem scoping, failure modes, metric frameworks, A/B or quasi-experiment setup, and post-launch monitoring without requiring implementation-level detail.
Community answers
Answer by SS
Here is a structured approach:1. Define Goals & Success Metrics
Primary Goal: Increase user engagement and expressiveness without significantly reducing text-based communication.
Key Metrics (KPIs):
Adoption Rate: % of Daily Active Users (DAU) who use a reaction within 7 days of launch.
Reaction Frequency: Total number of reactions sent per user/day.
Reaction per Message Rate: Ratio of reactions to total messages sent.
Reaction Diversity: Distribution of different emojis used (e.g., love, haha, wow, sad, angry).
Guardrail Metrics (What to Watch):
Message Volume/Chat Frequency: We must ensure reactions are additive to conversation, not replacing meaningful text messages.
Retention: D7/D30 retention of users who use the reaction feature vs. those who do not. [1, 2, 3, 4, 5]
Evaluation Framework (A/B Testing)I would propose an A/B test (experiment) comparing the new feature (Treatment) against the existing experience (Control). [1]
Treatment (A): Users can long-press a message to add an emoji reaction.
Control (B): Standard chat experience.
User Segmentation: Divide users by geographic region and usage volume (heavy vs. light chatters) to ensure unbiased, representative samples.
Duration: Run for at least two weeks to capture weekly behavior patterns. [1, 2, 3]
Detailed Analysis Plan
Adoption Funnel:
User receives a message.
User long-presses (Usage).
User selects an emoji (Selection).
Reaction appears in UI (Successful send).
Deep Dive on Engagement:
Long-press vs. Quick-reacti
Answer by SS
Here is a structured approach:1. Define Goals & Success Metrics
> Primary Goal:
Increase user engagement and expressiveness without significantly reducing text-based communication.
> Key Metrics (KPIs):
> Adoption Rate:
% of Daily Active Users (DAU) who use a reaction within 7 days of launch.
> Reaction Frequency:
Total number of reactions sent per user/day.
> Reaction per Message Rate:
Ratio of reactions to total messages sent.
> Reaction Diversity:
Distribution of different emojis used (e.g., love, haha, wow, sad, angry).
> Guardrail Metrics (What to Watch):
> Message Volume/Chat Frequency:
We must ensure reactions are additive to conversation, not replacing meaningful text messages.
> Retention:
D7/D30 retention of users who use the reaction feature vs. those who do not. [
1
,
2
,
3
,
4
,
5
]
Evaluation Framework (A/B Testing)I would propose an A/B test (experiment) comparing the new feature (Treatment) against the existing experience (Control). [
> 1
]
> Treatment (A):
Users can long-press a message to add an emoji reaction.
> Control (B):
Standard chat experience.
> User Segmentation:
Divide users by geographic region and usage volume (heavy vs. light chatters) to ensure unbiased, representative samples.
> Duration:
Run for at least two weeks to capture weekly behavior patterns. [
1
,
2
,
3
]
Detailed Analysis Plan
> Adoption Funnel:
User receives a message.
User long-presses (Usage).
User selects an emoji (Selection).
Reaction appears in UI (Successful send).
> Deep Div