Estimate experimental impact from cohort and timestamped metric data using user-level outcomes, difference in means, confidence intervals, and valid testing.
Estimate an Experiment’s Mean Effect from Cohort and Metric Tables
Company: LinkedIn
Role: Data Scientist
Category: Statistics & Math
Difficulty: medium
Interview Round: Onsite
# Estimate an Experiment’s Mean Effect from Cohort and Metric Tables
An experiment cohort table records each user, experiment-entry time, and treatment or control assignment. A metric table contains timestamped values for each user. Explain how you would define an outcome, assemble the analytical dataset, and estimate the treatment effect as a difference in means with a confidence interval. Also explain how a hypothesis test relates to that interval. The metric definition is left to you and must be stated explicitly; this is a statistical design discussion, not a query-writing task.
### What a Strong Answer Covers
- A prespecified user-level outcome and comparable observation window relative to assignment.
- A validated cohort join, complete follow-up, and correct treatment of missing observations.
- An effect estimate and uncertainty at the randomization unit rather than at the event-row level.
- Appropriate assumptions for inference and checks on allocation, dependence, and metric distributions.
```hint Aggregate before comparing
A user with many metric rows should not automatically count as many independently randomized people.
```
### Follow-up Questions
- When does an absent metric row mean zero rather than missing data?
- What changes if the metric is a ratio rather than a per-user mean?
Overview: Estimate experimental impact from cohort and timestamped metric data using user-level outcomes, difference in means, confidence intervals, and valid testing.
Estimate an Experiment’s Mean Effect from Cohort and Metric Tables
LinkedIn
Sep 22, 2026
mediumData ScientistOnsiteStatistics & Math
0
0
Estimate an Experiment’s Mean Effect from Cohort and Metric Tables
An experiment cohort table records each user, experiment-entry time, and treatment or control assignment. A metric table contains timestamped values for each user. Explain how you would define an outcome, assemble the analytical dataset, and estimate the treatment effect as a difference in means with a confidence interval. Also explain how a hypothesis test relates to that interval. The metric definition is left to you and must be stated explicitly; this is a statistical design discussion, not a query-writing task.
What a Strong Answer Covers Guidance
A prespecified user-level outcome and comparable observation window relative to assignment.
A validated cohort join, complete follow-up, and correct treatment of missing observations.
An effect estimate and uncertainty at the randomization unit rather than at the event-row level.
Appropriate assumptions for inference and checks on allocation, dependence, and metric distributions.
Follow-up Questions Guidance
When does an absent metric row mean zero rather than missing data?
What changes if the metric is a ratio rather than a per-user mean?