How would you evaluate stolen-post detection?

Quick Overview

This question evaluates product analytics, experimental design, and causal thinking for content-moderation algorithms—specifically metric specification, trade-off/harm analysis, and online experiment logistics—and is commonly asked to gauge a data scientist’s ability to balance detection accuracy, stakeholder impacts, and business objectives in production features; it is in the Analytics & Experimentation category for a Data Scientist position. At a high abstraction level it probes system-level reasoning around problem scoping, failure modes, metric frameworks, A/B or quasi-experiment setup, and post-launch monitoring without requiring implementation-level detail.

How would you evaluate stolen-post detection?

Company: Meta

Role: Data Scientist

Category: Analytics & Experimentation

Difficulty: hard

Interview Round: Technical Screen

You are interviewing for a Meta DSA (product analytics / data science) role. The product team is launching a new **Stolen Post Detection** algorithm that flags posts suspected of being copied/reposted without attribution, and then triggers actions (e.g., downrank, warning label, creator notification, or removal). Design an evaluation plan covering: 1) **Problem diagnosis & clarification:** What questions would you ask to clarify the product goal and the meaning of “stolen” (e.g., exact duplicate vs paraphrase vs meme templates), enforcement actions, and success criteria? 2) **Harms & tradeoffs:** Enumerate likely failure modes and harms of false positives vs false negatives, including different stakeholder impacts (original creator, reposter, viewers, moderators). 3) **Metrics:** Propose a metric framework with (a) primary success metrics, (b) guardrails, and (c) offline model metrics. Include at least one metric that can move in opposite directions depending on threshold choice. 4) **Experiment design:** Propose an online experiment (or quasi-experiment if A/B is hard). Address logging, unit of randomization, interference/network effects, ramp strategy, and how you would compute/think about power/MDE. 5) **Post-launch monitoring:** What would you monitor to detect regressions or gaming, and how would you iterate on thresholds/policy over time?

Overview: This question evaluates product analytics, experimental design, and causal thinking for content-moderation algorithms—specifically metric specification, trade-off/harm analysis, and online experiment logistics—and is commonly asked to gauge a data scientist’s ability to balance detection accuracy, stakeholder impacts, and business objectives in production features; it is in the Analytics & Experimentation category for a Data Scientist position. At a high abstraction level it probes system-level reasoning around problem scoping, failure modes, metric frameworks, A/B or quasi-experiment setup, and post-launch monitoring without requiring implementation-level detail.

Community answers

Answer by SS

Here is a structured approach:1. Define Goals & Success Metrics Primary Goal: Increase user engagement and expressiveness without significantly reducing text-based communication. Key Metrics (KPIs): Adoption Rate: % of Daily Active Users (DAU) who use a reaction within 7 days of launch. Reaction Frequency: Total number of reactions sent per user/day. Reaction per Message Rate: Ratio of reactions to total messages sent. Reaction Diversity: Distribution of different emojis used (e.g., love, haha, wow, sad, angry). Guardrail Metrics (What to Watch): Message Volume/Chat Frequency: We must ensure reactions are additive to conversation, not replacing meaningful text messages. Retention: D7/D30 retention of users who use the reaction feature vs. those who do not. [1, 2, 3, 4, 5] Evaluation Framework (A/B Testing)I would propose an A/B test (experiment) comparing the new feature (Treatment) against the existing experience (Control). [1] Treatment (A): Users can long-press a message to add an emoji reaction. Control (B): Standard chat experience. User Segmentation: Divide users by geographic region and usage volume (heavy vs. light chatters) to ensure unbiased, representative samples. Duration: Run for at least two weeks to capture weekly behavior patterns. [1, 2, 3] Detailed Analysis Plan Adoption Funnel: User receives a message. User long-presses (Usage). User selects an emoji (Selection). Reaction appears in UI (Successful send). Deep Dive on Engagement: Long-press vs. Quick-reacti

Answer by SS

Here is a structured approach:1. Define Goals & Success Metrics > Primary Goal: Increase user engagement and expressiveness without significantly reducing text-based communication. > Key Metrics (KPIs): > Adoption Rate: % of Daily Active Users (DAU) who use a reaction within 7 days of launch. > Reaction Frequency: Total number of reactions sent per user/day. > Reaction per Message Rate: Ratio of reactions to total messages sent. > Reaction Diversity: Distribution of different emojis used (e.g., love, haha, wow, sad, angry). > Guardrail Metrics (What to Watch): > Message Volume/Chat Frequency: We must ensure reactions are additive to conversation, not replacing meaningful text messages. > Retention: D7/D30 retention of users who use the reaction feature vs. those who do not. [ 1 , 2 , 3 , 4 , 5 ] Evaluation Framework (A/B Testing)I would propose an A/B test (experiment) comparing the new feature (Treatment) against the existing experience (Control). [ > 1 ] > Treatment (A): Users can long-press a message to add an emoji reaction. > Control (B): Standard chat experience. > User Segmentation: Divide users by geographic region and usage volume (heavy vs. light chatters) to ensure unbiased, representative samples. > Duration: Run for at least two weeks to capture weekly behavior patterns. [ 1 , 2 , 3 ] Detailed Analysis Plan > Adoption Funnel: User receives a message. User long-presses (Usage). User selects an emoji (Selection). Reaction appears in UI (Successful send). > Deep Div
|Home/Analytics & Experimentation/Meta
Meta logo
Meta
Mar 5, 2026
hardData ScientistTechnical ScreenAnalytics & Experimentation
112
0

You are interviewing for a Meta DSA (product analytics / data science) role. The product team is launching a new Stolen Post Detection algorithm that flags posts suspected of being copied/reposted without attribution, and then triggers actions (e.g., downrank, warning label, creator notification, or removal).

Design an evaluation plan covering:

  1. Problem diagnosis & clarification: What questions would you ask to clarify the product goal and the meaning of “stolen” (e.g., exact duplicate vs paraphrase vs meme templates), enforcement actions, and success criteria?
  2. Harms & tradeoffs: Enumerate likely failure modes and harms of false positives vs false negatives, including different stakeholder impacts (original creator, reposter, viewers, moderators).
  3. Metrics: Propose a metric framework with (a) primary success metrics, (b) guardrails, and (c) offline model metrics. Include at least one metric that can move in opposite directions depending on threshold choice.
  4. Experiment design: Propose an online experiment (or quasi-experiment if A/B is hard). Address logging, unit of randomization, interference/network effects, ramp strategy, and how you would compute/think about power/MDE.
  5. Post-launch monitoring: What would you monitor to detect regressions or gaming, and how would you iterate on thresholds/policy over time?
Loading comments...