Quantify launch decision with tests and guardrails

Read the full interview experience this question came from →

Quick Overview

This question evaluates a data scientist's competency in experimental design, statistical inference, and causal-effect estimation for clustered A/B tests, including sample size calculation with design effect, selection of cluster-robust hypothesis tests, sequential alpha-spending, multiple-testing control for guardrails, and quantifying contamination bias. It is commonly asked in the Statistics & Math domain to assess the ability to formalize decision rules that control type I/II errors and interpret confidence intervals under clustering and interim looks, testing both conceptual understanding of statistical principles and practical application of experiment governance.

Quantify launch decision with tests and guardrails

Company: Meta

Role: Data Scientist

Category: Statistics & Math

Difficulty: medium

Interview Round: Technical Screen

You will formalize the statistical decision rules for the Instagram button experiment described above. Given: baseline exploration rate (p0) = 0.15 per user-week, desired MDE = +1.5 percentage points (absolute), two-sided alpha = 0.05, power = 0.80. Randomization is at the cluster level with average cluster size m = 500 users and intra-cluster correlation ICC = 0.02. (a) Derive the required number of users and clusters per arm using a proportions test adjusted for clustering (design effect DE = 1 + (m − 1)·ICC). Show formulas and final numbers. (b) Specify the exact hypothesis test you would use for the primary metric (e.g., cluster-robust z-test on cluster means vs user-level test with cluster-robust standard errors). Explain when a nonparametric alternative would be preferable. (c) Define the confidence interval you will report and how you’ll interpret it jointly with practical significance. (d) You will look at the primary metric each week for 4 weeks. Choose and justify a sequential testing plan (e.g., O’Brien–Fleming alpha-spending), and show the adjusted per-look alphas. (e) You track 3 guardrails: p95 latency, crash rate, and add-to-cart rate. Describe a multiple-testing control that preserves power on the primary (e.g., hierarchical testing or Holm–Bonferroni) and write the decision logic combining primary and guardrails. (f) If contamination causes 10% of control users to see the button, quantify the bias direction for ITT and outline a correction (e.g., CACE with instrumented assignment).

Overview: This question evaluates a data scientist's competency in experimental design, statistical inference, and causal-effect estimation for clustered A/B tests, including sample size calculation with design effect, selection of cluster-robust hypothesis tests, sequential alpha-spending, multiple-testing control for guardrails, and quantifying contamination bias. It is commonly asked in the Statistics & Math domain to assess the ability to formalize decision rules that control type I/II errors and interpret confidence intervals under clustering and interim looks, testing both conceptual understanding of statistical principles and practical application of experiment governance.

Read the full Meta Data Scientist interview experience this question came from

Community answers

Answer by SS

(a) Given Baseline rate: ( p_0 = 0.15 ) MDE (absolute): ( \delta = 0.015 \Rightarrow p_1 = 0.165 ) Significance: ( \alpha = 0.05 ) (two-sided) → ( z_{1-\alpha/2} = 1.96 ) Power: ( 0.80 ) → ( z_{1-\beta} = 0.84 ) Cluster size: ( m = 500 ) ICC: ( \rho = 0.02 ) Step 1: Sample size (ignoring clustering) Variance = 0.2653 MDE = 0.015 Users per arm = 7.86 0.2653/0.0150.015 Users per arm (no clustering): ~9,244 Step 2: Adjust for clustering (Design Effect) DE = 1 + (500- 1)* 0.02 = 1 + 499 *0.02 = 1 + 9.98 = 10.98 Step 3: Effective sample size with clustering Users when clustering = 10.98 * 9244 Users per arm (adjusted): ~101,500 Step 4: Convert users → clusters 101,500/500 = 203 Final Answer Per arm: Users required: ~101,500 Clusters required: ~203 Total experiment: Users: ~203,000 Clusters: ~406 Key Insight The design effect ≈ 11× dramatically increases sample size due to clustering. Even a modest ICC (0.02) becomes very costly when cluster size is large (500).
|Home/Statistics & Math/Meta
Meta logo
Meta
Oct 13, 2025
mediumData ScientistTechnical ScreenStatistics & Math
4
0

You will formalize the statistical decision rules for the Instagram button experiment described above. Given: baseline exploration rate (p0) = 0.15 per user-week, desired MDE = +1.5 percentage points (absolute), two-sided alpha = 0.05, power = 0.80. Randomization is at the cluster level with average cluster size m = 500 users and intra-cluster correlation ICC = 0.02. (a) Derive the required number of users and clusters per arm using a proportions test adjusted for clustering (design effect DE = 1 + (m − 1)·ICC). Show formulas and final numbers. (b) Specify the exact hypothesis test you would use for the primary metric (e.g., cluster-robust z-test on cluster means vs user-level test with cluster-robust standard errors). Explain when a nonparametric alternative would be preferable. (c) Define the confidence interval you will report and how you’ll interpret it jointly with practical significance. (d) You will look at the primary metric each week for 4 weeks. Choose and justify a sequential testing plan (e.g., O’Brien–Fleming alpha-spending), and show the adjusted per-look alphas. (e) You track 3 guardrails: p95 latency, crash rate, and add-to-cart rate. Describe a multiple-testing control that preserves power on the primary (e.g., hierarchical testing or Holm–Bonferroni) and write the decision logic combining primary and guardrails. (f) If contamination causes 10% of control users to see the button, quantify the bias direction for ITT and outline a correction (e.g., CACE with instrumented assignment).

Loading comments...