Diagnose Discrepancy in A/B Test Conversion Rate Results
Quick Overview
Evaluates A/B test design and discrepancy diagnosis for personalized marketing emails. Strong answers define user-level randomization, metrics, power, and guardrails, then explain why a later rollout might show smaller lift through population, seasonality, implementation, instrumentation, or statistical issues.
Diagnose Discrepancy in A/B Test Conversion Rate Results
Company: Coinbase
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Onsite
##### Scenario
An e-commerce firm plans to send personalized marketing emails to increase purchase conversions and wants to rigorously evaluate the impact.
##### Question
Design an A/B test to measure whether personalized emails lift conversion rate. Specify the primary and guardrail metrics, statistical test, minimum detectable effect, required sample size and duration. After launching on the full user base, a new director reruns the test and sees only a 2 % lift versus the original 20 %. List possible causes and the analyses you would run to diagnose the discrepancy.
##### Hints
Think through experiment design, power analysis, instrumentation issues, seasonality, user overlap, novelty effects and segmentation cuts.
Overview: Evaluates A/B test design and discrepancy diagnosis for personalized marketing emails. Strong answers define user-level randomization, metrics, power, and guardrails, then explain why a later rollout might show smaller lift through population, seasonality, implementation, instrumentation, or statistical issues.
Community answers
Answer by Jay123
for 20% vs 2%, Here's what usually goes wrong when a test that worked before suddenly doesn't:
1. Different audience
Maybe the first test hit your best customers (cart abandoners, recent visitors) but the rerun went to everyone. Check if the audience mix changed and run the numbers by segment.
2. Timing matters
First test during Black Friday, rerun in dead January? That'll do it. Compare baseline conversion rates between periods and consider seasonality.
3. Novelty wore off
People loved the personalization at first, now they're used to it. Track lift by how many emails they've gotten and how long since their first personalized one.
4. Email deliverability tanked
Scaling up hurt your sender reputation. Check if emails are hitting spam folders more often - you'll see the drop at the delivered→opened stage, not clicks.
5. Metrics changed
Someone tweaked attribution windows or event definitions between tests. Rerun both tests with identical definitions to be sure.
6. Control group contamination
Your "control" users got similar emails from other campaigns. Check if they received any other marketing emails during the test window.
7. Other channels interfering
SMS blast went to treatment but not control? That's a problem. Check what other marketing touched each group.
8. Read results too early
Someone got impatient and called it at day 2 instead of waiting the full week. Make sure attribution windows actually closed.
9. Randomization broke
Groups aren't actually random. Run an SRM t
Diagnose Discrepancy in A/B Test Conversion Rate Results
An e-commerce company plans to send personalized marketing emails to increase purchase conversions. An initial experiment showed a large lift, but after broader rollout a later test found a much smaller lift.
Constraints & Assumptions
Focus on rigorous experiment design and post-hoc diagnosis, not on building the personalization model itself.
Assume user-level email eligibility, send, open, click, purchase, unsubscribe, and revenue data are available.
Conversion lift should be measured causally against a control group.
Consider both statistical explanations and real product or implementation explanations.
Clarifying Questions to Ask Guidance
What was the original population, and how did it differ from the later rollout population?
Were send time, subject line, cadence, offer, and creative held constant?
Was assignment persistent at the user level?
Was the 20% lift relative or absolute, and over what conversion window?
Part 1 - Design the A/B Test
Design an A/B test to measure whether personalized emails increase conversion rate.
What This Part Should Cover Guidance
Unit of randomization, exposure rules, control and treatment definitions, and eligibility criteria.
Primary metric such as purchase conversion, plus secondary and guardrail metrics like revenue, unsubscribes, spam complaints, margin, and long-term retention.
Instrumentation checks, sample-ratio mismatch checks, and pre-registration of success criteria.
Part 2 - Diagnose the Lift Discrepancy
After full rollout, a new director reruns the test and observes only a 2% lift instead of the original 20%. List plausible causes and the analyses you would run.
What This Part Should Cover Guidance
Differences in population, seasonality, campaign creative, offer, email deliverability, model version, product changes, and competitor or market conditions.
Statistical issues such as underpowered tests, novelty effects, peeking, multiple testing, regression to the mean, or sample-ratio mismatch.
Implementation problems such as treatment contamination, incorrect logging, duplicate sends, personalization not actually applied, or inconsistent attribution windows.
Segment analysis and reanalysis using the original experiment definition where possible.
What a Strong Answer Covers Guidance
A strong answer designs a clean user-level experiment, quantifies power and launch criteria, and diagnoses the later discrepancy with specific checks rather than vague speculation.
Follow-up Questions Guidance
How would you design a long-term holdout after rollout?
What if open rate increases but purchase conversion does not?
How would you explain relative versus absolute lift to a non-technical stakeholder?