Master A/B Testing: Key Concepts and Methodologies Explained

Quick Overview

Evaluates practical A/B testing and causal inference fundamentals for online products. Strong answers cover p-values, errors, power, workflow, Simpson's paradox, metrics, and observational causal methods.

Master A/B Testing: Key Concepts and Methodologies Explained

Company: PayPal

Role: Data Scientist

Category: Analytics & Experimentation

Difficulty: medium

Interview Round: Onsite

##### Scenario Data scientist is interviewed on A/B-testing know-how for an online product. ##### Question Explain what a p-value represents; define Type I and Type II errors; outline the end-to-end experimentation workflow; describe Simpson's paradox and how to detect it; propose primary/secondary metrics; name two causal-inference methods useful when randomization is impossible and when you would apply them. ##### Hints Cover hypothesis, sample-size, segmentation, lift vs variance, DAGs or matching, and practical examples.

Overview: Evaluates practical A/B testing and causal inference fundamentals for online products. Strong answers cover p-values, errors, power, workflow, Simpson's paradox, metrics, and observational causal methods.

|Home/Analytics & Experimentation/PayPal
PayPal logo
PayPal
Jul 12, 2025
mediumData ScientistOnsiteAnalytics & Experimentation
80
0

A/B Testing and Causal Inference: Core Concepts

You are a data scientist interviewing for a role working on an online product. Demonstrate practical A/B testing and causal inference knowledge.

Provide concise, accurate explanations and guidance for the topics below.

Constraints & Assumptions

  • Use practical product examples, not only definitions.
  • Distinguish statistical significance, practical significance, and causal validity.
  • Include experiment design, metric choice, variance, segmentation, and decision-making.
  • Mention causal inference only with the assumptions required.

Clarifying Questions to Ask Guidance

  • What type of online product and metric are we experimenting on?
  • Is the question about randomized A/B testing or observational causal inference?
  • What business risk is associated with false positives and false negatives?
  • Are there network effects, interference, or delayed outcomes?

Part 1 - P-values, Errors, and Power

Explain p-values, Type I and Type II errors, and power.

What This Part Should Cover Guidance

  • Define a p-value under the null hypothesis and state common misinterpretations.
  • Define Type I error, Type II error, alpha, beta, and power.
  • Explain how sample size, variance, effect size, alpha, and test duration interact.

Part 2 - Experimentation Workflow

Outline an end-to-end workflow from hypothesis to decision.

What This Part Should Cover Guidance

  • Define hypothesis, randomization unit, exposure, eligibility, metrics, guardrails, sample size, duration, and analysis plan.
  • Include SRM, logging checks, pre-period balance, CUPED or variance reduction, segmentation, and novelty or seasonality.
  • Decide using pre-defined criteria and practical impact.

Part 3 - Simpson's Paradox and Metrics

Define Simpson's paradox and propose primary and secondary metrics for an online product experiment.

What This Part Should Cover Guidance

  • Explain aggregate versus stratified result reversals caused by confounding or imbalance.
  • Give a practical example and how to detect or handle it.
  • Choose primary, secondary, diagnostic, and guardrail metrics with clear definitions.

Part 4 - Causal Inference Beyond A/B Tests

Explain when and how you would use causal inference methods outside randomized experiments.

What This Part Should Cover Guidance

  • Mention matching, difference-in-differences, regression discontinuity, instrumental variables, synthetic control, or inverse-propensity weighting where appropriate.
  • State assumptions such as ignorability, overlap, parallel trends, exclusion restriction, or continuity.
  • Validate assumptions and report uncertainty.

Follow-up Questions Guidance

  • What would you do if an experiment has sample ratio mismatch?
  • How would you analyze treatment effects that differ by user segment?
  • When would you refuse to make a causal claim from observational data?
Loading comments...