Analyze Distribution of Daily Page Shares Per User

Quick Overview

Meta statistics prompt on daily page-share and time-spent distributions, covering zero inflation, heavy tails, percentiles, cohort regression, negative-binomial-style models, variance, and tail stability.

Analyze Distribution of Daily Page Shares Per User

Company: Meta

Role: Data Scientist

Category: Statistics & Math

Difficulty: medium

Interview Round: Onsite

##### Question Sketch the distribution of daily page shares per user, indicating mean, median, p1, and p99. What shape do you expect and why? For users at the 50th and 95th percentiles of today’s distribution, predict their average shares two weeks from now. For the cohort with exactly two shares on day 1, describe the expected trend of average shares over days 2–30. Repeat for the cohort with five shares; which cohort will have larger variance and what distribution do you expect? Repeat the above style of analysis for daily time-spent-per-user; comment on stability of mean and tail behavior over three weeks. ##### Hints Engagement metrics are typically heavy-tailed; cohorts regress toward the overall mean while retaining right-skewed structure.

Overview: Meta statistics prompt on daily page-share and time-spent distributions, covering zero inflation, heavy tails, percentiles, cohort regression, negative-binomial-style models, variance, and tail stability.

Community answers

Answer by SS

(2) The core concept: regression to the mean Before even thinking about regression models, ask: what happens naturally to extreme values over time? The p50 user (0 shares today) is a typical user. They'll stay roughly typical. Not much change expected. The p95 user (say, 8 shares today) is unusually active today. But are they genuinely a high-sharer, or did they just have an unusually active day? Likely some of both. Over the next 14 days, the extreme behavior will partially "revert" back toward their true average. This is called regression to the mean — extreme observations today are partly signal (their true rate) and partly noise (a lucky/unlucky day). So directionally: p50 user → stays near 0, maybe slightly above p95 user → comes down noticeably from their day-1 peak, settles at something lower Why simple linear regression isn't quite right here Linear regression would work if the outcome were continuous and symmetric. But daily shares are: Count data (0, 1, 2, 3...), not continuous Zero-inflated and right-skewed Non-negative by definition Predicting negative shares is nonsensical, but linear regression doesn't know that. What you'd actually use A Negative Binomial regression (since we already established shares are overdispersed count data) where: Outcome: avg daily shares over the next 14 days Key predictor: today's share count (or better, their historical average) Other predictors you'd add: days active on platform, avg shares over past 30 days, user tenure The histor

Answer by SS

(3) Cohort with 2 shares on day 1 2 shares is slightly above the median (which is 0) but not extreme. These users are mildly active. Their day 1 value is a mix of: Some genuine mild-sharers (true signal) Some typical 0-share users who just happened to share twice (noise) Over days 2–30, the noise users revert to 0, dragging the cohort average down. The genuine mild-sharers stay around 1–2. So the trend looks like: Starts at 2 → drops in the first few days → stabilizes at some value below 2, maybe around 0.8–1.2 The drop is moderate and the stabilization happens relatively quickly. Cohort with 5 shares on day 1 5 shares is quite extreme — well into the right tail. This cohort is a mix of: A few genuine power users (true signal) Many moderate users who had an unusually active day (noise) The noise component is much larger here because 5 shares is so far from the typical user's behavior. So over days 2–30: Starts at 5 → drops sharply in the first few days → stabilizes much lower, maybe around 1.5–2.5 The drop is steeper and more dramatic than the 2-share cohort. This is regression to the mean hitting harder because the starting point was more extreme. Visualizing the trend--- Which cohort has larger variance? The 5-share cohort, and by a lot. Here's why: The 5-share cohort is a much more mixed bag of people. It contains: Genuine power users who'll stay at 5–8 shares/day Moderate users who'll drop to 1–2 Typical users who had a freak day and will drop to 0 That wide spread of und
|Home/Statistics & Math/Meta
Meta logo
Meta
Jul 12, 2025
mediumData ScientistOnsiteStatistics & Math
88
0

Engagement Distributions and Cohort Dynamics

You are analyzing per-user, per-day engagement. Assume the panel includes all users, inactive days count as zeros, and bots or obvious spam accounts have been removed.

Constraints & Assumptions

  • Daily page shares are count data and are likely zero-inflated, overdispersed, and right-skewed.
  • Users have heterogeneous long-run sharing propensities and day-to-day randomness.
  • Percentile-selected cohorts can regress toward the population mean over time.
  • Daily time spent is nonnegative, heavy-tailed, and may be censored or capped by measurement rules.

Clarifying Questions to Ask Guidance

  • Are inactive users included as zero-share days?
  • Are we analyzing users, active users, or sessions?
  • Is the goal descriptive reporting, forecasting, or anomaly detection?
  • Are shares and time spent measured consistently across platforms and time zones?

What a Strong Answer Covers Guidance

  • Distribution sketch for daily shares: a spike at zero, long right tail, mean greater than median, p1 often zero, and p99 far to the right.
  • Explanation of zero inflation, heavy tails, heterogeneous user propensities, burstiness, and event-driven behavior.
  • Prediction for users at today's 50th and 95th percentiles: both should be estimated from historical transition/cohort data, with high-percentile users regressing downward while remaining above average.
  • Cohort trajectories for exactly 2 and exactly 5 shares on day 1, including regression toward each cohort's latent mean and wider variance for the higher-activity cohort.
  • Suitable count distributions or models, such as zero-inflated negative binomial, Poisson-gamma mixtures, hurdle models, or user-level random effects.
  • Similar analysis for time spent, including heavy tail, stability of mean versus tail percentiles, winsorization/capping, and cohort persistence over three weeks.

Follow-up Questions Guidance

  • Why can the mean be unstable for heavy-tailed engagement metrics?
  • How would you estimate the two-week forecast empirically?
  • What would you do if p99 jumps but the median is unchanged?
  • How would bot filtering change the distribution?
Loading comments...