A/B Testing And Experiment Design
Asked of: Data Scientist
Last updated

What's being tested
Candidates must demonstrate end-to-end mastery of randomized experiment design and analysis for product decisions: choosing the correct unit of randomization, defining primary and guardrail metrics, computing sample sizes and power, and producing a pre-registered analysis that handles clustering, covariate adjustment, and possible interference. Interviewers probe statistical reasoning (Type I/II tradeoffs, sequential testing), practical variance-reduction techniques (e.g., CUPED), and operational diagnostics for ambiguous or null results. The goal is to show you can deliver a causal, defensible answer a product team can act on.
Core knowledge
-
Randomization unit: choose between user, session, ad-impression, or seller based on treatment scope and interference; wrong unit creates contamination or underpowered experiments.
-
Hypotheses & estimand: state null/alternative and exact estimand (e.g., average treatment effect on users who see the new algorithm), not just “does it improve CTR”.
-
Sample-size formula: for two-sample mean test, per-group n = ((z_{1-α/2}+z_{1-β})^2 * 2σ^2) / Δ^2; estimate σ from pre-experiment data and set minimum detectable effect Δ (SESOI).
-
Power tradeoffs: doubling sample reduces detectable Δ by √2; longer duration vs. larger population—account for seasonal/weekday patterns when estimating duration.
-
Cluster randomization & ICC: design effect = 1 + (m−1)*ICC inflates required sample when randomizing clusters; estimate intraclass correlation (ICC) from historical data.
-
Covariate adjustment / CUPED: subtract linear projection on baseline covariate to reduce variance; works best when pre-period metric correlates strongly with outcome (use regression-based adjustment in
R/Python). -
Sequential testing & peeking: avoid naive peeking; use alpha-spending (Pocock/O’Brien–Fleming) or validated sequential methods (e.g., always-valid p-values) to preserve Type I error.
-
Multiple comparisons: correct for familywise error when testing many variants/metrics (Bonferroni, Holm) or control FDR (Benjamini–Hochberg) and predefine metric hierarchy (primary vs. guardrails).
-
Interference & exposure models: if SUTVA violated, define exposure mapping (who is exposed to treated peers) and consider cluster assignment, two-stage randomization, or partial population designs.
-
Analysis plan: pre-specify estimand, test statistic, transformation (log, winsorize), outlier handling, and whether to use cluster-robust SE, permutation tests, or bootstrapping for inference.
-
Guardrails and safety metrics: always include business guardrails (e.g., revenue, retention, quality signals) and health metrics (latency, error rate); use sequential monitoring but mute decisions until final.
-
Null results interpretation: distinguish low power from true null with confidence intervals, equivalence testing, and reporting of minimal detectable effects; quantify uncertainty for product decisions.
-
Diagnostics: randomization checks, instrumentation loss, skewed traffic splits, and differential attrition; check balance on pre-period metrics and stratify or reweight if needed.
Worked example — "Design an A/B test for a new shop-ads algorithm"
First 30 seconds: ask clarifying questions — what’s the treatment (ranking model change?), scope (all shoppers or subset?), cost of misassignment, and primary business metric (e.g., ad revenue per DAU vs. purchase conversion). State assumptions: stable traffic, instrumentation available via BigQuery logs, pre-period metric for CUPED. Organize the answer into pillars: (1) unit of randomization — randomize at user to avoid the same shopper seeing mixed algorithms (unless seller-level constraints force clustering), (2) metrics — primary: revenue per user (log-transform), guardrails: CTR, purchase-rate, organic engagement, (3) power/sample-size — compute n using historical σ and SESOI, inflate for expected ICC or baseline variability, (4) analysis plan — pre-register intent-to-treat ATE, use CUPED with pre-period revenue, apply cluster-robust SE if clusters used, (5) monitoring & roll-out — limited ramp, abort thresholds tied to guardrails. Flag a tradeoff: randomizing at seller vs. user trades statistical efficiency for reduced interference; state rule-of-thumb: prefer user-level unless seller-level interference is large. Close with next steps: if time allowed, plan heterogeneous treatment effect analysis by cohort, long-term retention measurement, and an uplift model to personalize decisions.
A second angle — "Design and analyze an A/B test with interference"
Interference forces you to redefine the estimand: instead of ATE assume partial interference and specify spillover estimands (direct vs. indirect effects). Use cluster randomization (e.g., households, social-graph clusters) or two-stage randomization to estimate spillovers explicitly. If clusters are impractical, build an exposure model mapping which users are plausibly affected and analyze conditional ATEs (those exposed vs. not). Analysis-wise, rely on permutation tests respecting cluster boundaries or use randomization inference to get valid p-values under interference. The same fundamentals (power, pre-specification, guardrails) apply, but sample-size must account for between-cluster variance and the loss of effective sample due to contamination.
Common pitfalls
Pitfall: Ignoring clustering. A tempting quick-sample calculation assuming i.i.d. users will understate required sample when ICC>0; always estimate ICC and apply the design effect to sample-size.
Pitfall: Over-interpreting p-values from a peeked experiment. Early peeking without sequential corrections yields inflated Type I error; report always-valid intervals or use pre-defined stopping rules.
Pitfall: Reporting a null as “no effect” without SESOI. A non-significant result might be underpowered; present the confidence interval and whether it excludes the smallest effect size of interest.
Connections
Interviewers may pivot to heterogeneous treatment effect estimation (uplift models) or to observational causal inference methods when randomization isn’t feasible. They may also ask about experiment-velocity infrastructure (monitoring plans, but not pipeline internals), or ML model evaluation for ranking metrics.
Further reading
-
Kohavi et al., "Online Controlled Experiments at Large Scale" — practical lessons from industry on design and pitfalls.
-
Deng, Lu, and Kohavi, "Improving Online Controlled Experiments with Variance Reduction" (CUPED) — technique and empirical guidance.
-
Hudgens & Halloran, "Toward Causal Inference With Interference" — formal treatment of interference and estimands.
Featured in interview prep guides
Practice questions
- How to evaluate a new homepage featurePayPal · Data Scientist · Technical Screen · easy
- Design and Test a New FeatureUber · Data Scientist · Onsite · easy
- Design metrics and experiment for donation featurePayPal · Data Scientist · Onsite · easy
- Analyze an A/B test and present recommendationInstacart · Data Scientist · Onsite · medium
- Design and analyze a group-calls experimentMeta · Data Scientist · Technical Screen · medium
- Define composite success for search and test itMeta · Data Scientist · Onsite · medium
- Design B2C chatbot success metrics and test planMeta · Data Scientist · Onsite · medium
- Explain and validate A/B test assumptionsUber · Data Scientist · Technical Screen · hard
- Evaluate Facebook Dating launch and validate successMeta · Data Scientist · Technical Screen · hard
- Design and validate ad model launchMeta · Data Scientist · Onsite · medium
- Design and power an incentive experimentUber · Data Scientist · Onsite · hard
- Design an A/B test for promo-targeting modelsUber · Data Scientist · Technical Screen · hard
Related concepts
- A/B Testing And Experiment DesignAnalytics & Experimentation
- A/B Testing And Experiment DesignAnalytics & Experimentation
- A/B Testing, Power, And Experiment DesignAnalytics & Experimentation
- A/B Testing And Causal ExperimentationAnalytics & Experimentation
- A/B TestingAnalytics & Experimentation
- A/B TestingAnalytics & Experimentation