Analyze Distribution of Daily Page Shares Per User
Quick Overview
Meta statistics prompt on daily page-share and time-spent distributions, covering zero inflation, heavy tails, percentiles, cohort regression, negative-binomial-style models, variance, and tail stability.
Analyze Distribution of Daily Page Shares Per User
Company: Meta
Role: Data Scientist
Category: Statistics & Math
Difficulty: medium
Interview Round: Onsite
##### Question
Sketch the distribution of daily page shares per user, indicating mean, median, p1, and p99. What shape do you expect and why? For users at the 50th and 95th percentiles of today’s distribution, predict their average shares two weeks from now. For the cohort with exactly two shares on day 1, describe the expected trend of average shares over days 2–30. Repeat for the cohort with five shares; which cohort will have larger variance and what distribution do you expect? Repeat the above style of analysis for daily time-spent-per-user; comment on stability of mean and tail behavior over three weeks.
##### Hints
Engagement metrics are typically heavy-tailed; cohorts regress toward the overall mean while retaining right-skewed structure.
Overview: Meta statistics prompt on daily page-share and time-spent distributions, covering zero inflation, heavy tails, percentiles, cohort regression, negative-binomial-style models, variance, and tail stability.
Community answers
Answer by SS
(2)
The core concept: regression to the mean
Before even thinking about regression models, ask: what happens naturally to extreme values over time?
The p50 user (0 shares today) is a typical user. They'll stay roughly typical. Not much change expected.
The p95 user (say, 8 shares today) is unusually active today. But are they genuinely a high-sharer, or did they just have an unusually active day?
Likely some of both. Over the next 14 days, the extreme behavior will partially "revert" back toward their true average. This is called regression to the mean — extreme observations today are partly signal (their true rate) and partly noise (a lucky/unlucky day).
So directionally:
p50 user → stays near 0, maybe slightly above
p95 user → comes down noticeably from their day-1 peak, settles at something lower
Why simple linear regression isn't quite right here
Linear regression would work if the outcome were continuous and symmetric. But daily shares are:
Count data (0, 1, 2, 3...), not continuous
Zero-inflated and right-skewed
Non-negative by definition
Predicting negative shares is nonsensical, but linear regression doesn't know that.
What you'd actually use
A Negative Binomial regression (since we already established shares are overdispersed count data) where:
Outcome: avg daily shares over the next 14 days
Key predictor: today's share count (or better, their historical average)
Other predictors you'd add: days active on platform, avg shares over past 30 days, user tenure
The histor
Answer by SS
(3)
Cohort with 2 shares on day 1
2 shares is slightly above the median (which is 0) but not extreme. These users are mildly active. Their day 1 value is a mix of:
Some genuine mild-sharers (true signal)
Some typical 0-share users who just happened to share twice (noise)
Over days 2–30, the noise users revert to 0, dragging the cohort average down. The genuine mild-sharers stay around 1–2. So the trend looks like:
Starts at 2 → drops in the first few days → stabilizes at some value below 2, maybe around 0.8–1.2
The drop is moderate and the stabilization happens relatively quickly.
Cohort with 5 shares on day 1
5 shares is quite extreme — well into the right tail. This cohort is a mix of:
A few genuine power users (true signal)
Many moderate users who had an unusually active day (noise)
The noise component is much larger here because 5 shares is so far from the typical user's behavior. So over days 2–30:
Starts at 5 → drops sharply in the first few days → stabilizes much lower, maybe around 1.5–2.5
The drop is steeper and more dramatic than the 2-share cohort. This is regression to the mean hitting harder because the starting point was more extreme.
Visualizing the trend---
Which cohort has larger variance?
The 5-share cohort, and by a lot. Here's why:
The 5-share cohort is a much more mixed bag of people. It contains:
Genuine power users who'll stay at 5–8 shares/day
Moderate users who'll drop to 1–2
Typical users who had a freak day and will drop to 0
That wide spread of und
Analyze Distribution of Daily Page Shares Per User
Meta
Jul 12, 2025
mediumData ScientistOnsiteStatistics & Math
88
0
Engagement Distributions and Cohort Dynamics
You are analyzing per-user, per-day engagement. Assume the panel includes all users, inactive days count as zeros, and bots or obvious spam accounts have been removed.
Constraints & Assumptions
Daily page shares are count data and are likely zero-inflated, overdispersed, and right-skewed.
Users have heterogeneous long-run sharing propensities and day-to-day randomness.
Percentile-selected cohorts can regress toward the population mean over time.
Daily time spent is nonnegative, heavy-tailed, and may be censored or capped by measurement rules.
Clarifying Questions to Ask Guidance
Are inactive users included as zero-share days?
Are we analyzing users, active users, or sessions?
Is the goal descriptive reporting, forecasting, or anomaly detection?
Are shares and time spent measured consistently across platforms and time zones?
What a Strong Answer Covers Guidance
Distribution sketch for daily shares: a spike at zero, long right tail, mean greater than median, p1 often zero, and p99 far to the right.
Explanation of zero inflation, heavy tails, heterogeneous user propensities, burstiness, and event-driven behavior.
Prediction for users at today's 50th and 95th percentiles: both should be estimated from historical transition/cohort data, with high-percentile users regressing downward while remaining above average.
Cohort trajectories for exactly 2 and exactly 5 shares on day 1, including regression toward each cohort's latent mean and wider variance for the higher-activity cohort.
Suitable count distributions or models, such as zero-inflated negative binomial, Poisson-gamma mixtures, hurdle models, or user-level random effects.
Similar analysis for time spent, including heavy tail, stability of mean versus tail percentiles, winsorization/capping, and cohort persistence over three weeks.
Follow-up Questions Guidance
Why can the mean be unstable for heavy-tailed engagement metrics?
How would you estimate the two-week forecast empirically?
What would you do if p99 jumps but the median is unchanged?