Design Experiment to Test New Hashtag Recommender Algorithm

Quick Overview

Evaluates online experiment design for a new hashtag recommender in a social app. Strong answers define randomization, stratification, power, guardrails, interference, ramp, and analysis plan.

Design Experiment to Test New Hashtag Recommender Algorithm

Company: Meta

Role: Data Scientist

Category: Analytics & Experimentation

Difficulty: medium

Interview Round: Onsite

##### Scenario Evaluating a new hashtag recommendation algorithm through experimentation ##### Question How would you design and run an experiment to test the new hashtag recommender, specifically how would you choose test and control groups? ##### Hints Cover randomization unit, stratification, sample size, guardrail metrics, exposure overlap, and statistical power.

Quick Answer: Evaluates online experiment design for a new hashtag recommender in a social app. Strong answers define randomization, stratification, power, guardrails, interference, ramp, and analysis plan.

|Home/Analytics & Experimentation/Meta
Meta logo
Meta
Jul 12, 2025, 6:59 PM
mediumData ScientistOnsiteAnalytics & Experimentation
13
0

Experiment Design: Testing a New Hashtag Recommender

A social app shows hashtag recommendations to users while they compose posts. A new algorithm is proposed to increase hashtag adoption and downstream engagement without harming reliability, content quality, or user trust.

Design an online experiment to evaluate the new recommender.

Constraints & Assumptions

  • Define the triggered population and exposure clearly.
  • Account for overlap and interference because hashtags can affect downstream content distribution.
  • Include adoption, engagement, quality, and safety guardrails.
  • Predefine ramp, duration, power, and analysis plan.

Clarifying Questions to Ask Guidance

  • Who is eligible to see recommendations, and when are recommendations triggered?
  • What is the primary objective: acceptance, engagement, creator retention, or content quality?
  • Are creators or viewers the analysis unit for the primary metric?
  • Can hashtag trends create spillovers between treatment and control users?

Part 1 - Test and Control Groups

Choose the randomization unit, treatment/control definition, and exposure logging.

What This Part Should Cover Guidance

  • Prefer stable creator-level randomization for composer recommendations unless interference requires clusters.
  • Define control as the existing recommender and treatment as the new algorithm.
  • Log triggered exposure, recommended hashtags, positions, accepted tags, edits, and post outcomes.
  • Consider stratification by creator activity, region, language, platform, and content type.

Part 2 - Metrics, Power, and Run Plan

Define sample size, statistical power, guardrails, ramp, and duration.

What This Part Should Cover Guidance

  • Include primary metrics such as hashtag acceptance rate, posts with accepted hashtags, or engagement lift.
  • Include guardrails for post quality, hides/reports, spam, latency, crashes, creator retention, and viewer experience.
  • Estimate MDE, baseline rate, variance, power, alpha, and maturation window.
  • Use staged ramps, SRM checks, logging checks, and novelty monitoring.

Part 3 - Interference and Analysis

Address exposure overlap, interference, and final analysis.

What This Part Should Cover Guidance

  • Discuss feed-level spillovers, trending hashtags, creator-viewer interactions, and network effects.
  • Consider cluster randomization, geo/language holdouts, or switchbacks if spillovers are material.
  • Analyze intent-to-treat and triggered users carefully.
  • Segment by creator type, content type, language, and baseline activity.

Follow-up Questions Guidance

  • What if the new recommender increases adoption but also increases spam reports?
  • How would you handle creators who see recommendations on multiple devices?
  • How would you evaluate long-term hashtag ecosystem health?
Loading comments...