As a Data Scientist at Snap, you sit at the intersection of quantitative rigor and high-velocity product innovation. You drive the analytical engine behind core ecosystem products like Snapchat, Lens Studio, and Spectacles, translating massive, complex datasets into clear, actionable business strategies. Your work directly shapes how millions of people globally communicate, express themselves, and interact with augmented reality.
This role requires you to act as both a truth-seeker and a strategic advisor. Whether you are optimizing monetization funnels, scaling growth loops, or evaluating new messaging features, your insights steer roadmap prioritization across engineering, product management, and design. You will tackle ambiguous, high-impact problems where defining the right metric is just as important as building the predictive model behind it.
Expect an environment characterized by rapid experimentation and deep technical ownership. You will not just pull numbers; you will design rigorous A/B tests, diagnose unexpected metric fluctuations, and deploy scalable statistical solutions. If you thrive on ambiguity, possess strong product intuition, and want your analysis to directly influence a globally recognized product, this role offers an ideal platform.
Recruiter Conversation
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Screen
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Virtual Onsite Loop
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub editorial advice for the preparation topics above.
Crediting notifications with the sessions that follow them
Members who open a push notification were already more likely to open the app, so attributing their session to the notification measures intent rather than causation and makes almost any send look profitable. The predictable result is a push-volume increase that shows a large modelled gain and a small real one, paid for later in opt-outs and uninstalls that no single experiment window captures. The only defensible estimate compares a stable send arm against a holdout arm assigned at the decision point, with the held-out decisions logged and suppressed at delivery, over a window long enough to include the opt-out response. Frequency effects are non-linear, so a per-notification incremental rate estimated at one volume does not extrapolate to a higher one.
Using report volume as a measurement of how much violating content exists
Reporting is a member behaviour, not an observation of the content. Report counts rise when the report control is made easier to reach, when a coordinated campaign targets an account, and when the audience shifts toward people who object; they fall when violating content is shown mainly to members who agree with it. A ranker that gets better at matching bad content to receptive audiences will drive reports down and harm up at the same time. Prevalence must come from a random sample of served impressions with recorded selection probabilities, labelled by humans against the written policy, and reported with an interval. Reports are useful as a detection signal and as a demand-side complaint rate, not as a denominator-anchored measure of harm.
Answering a product-sense question with a list of features
Answer with a decision and the measurement that would settle it: the hypothesis, the primary metric, the guardrails, and the result that would make you not ship. A feature brainstorm cannot be wrong, which is exactly why it earns no points.
Reading an observational correlation as a causal effect
Name the confounder you are most worried about and the design that would remove it: an experiment, a difference-in-differences with a checked pre-period trend, an instrument, or a regression discontinuity. When none is available, state which direction the bias likely runs and bound the claim accordingly.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you test for heterogeneous treatment effects in an offline data…
How do you test for heterogeneous treatment effects in an offline data sample?
Approach
- Sanity-check the answer against a simple bound or a simulated case.
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
Explain the core assumptions behind linear regression and how you iden…
Explain the core assumptions behind linear regression and how you identify violations in observational data.
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Off-policy estimate with a positivity and weight audit
logged has impression_id, rank_position, reward (1 if the viewer engaged), log_propensity (float, NULL where serving was deterministic top-k) and target_propensity (the candidate policy's probability of placing the same item in the same slot, precomputed). Estimate the candidate policy's engagement rate per impression with inverse propensity scoring and with self-normalised IPS. Report the effective sample size as the square of the sum of weights over the sum of squared weights, the 99th percentile weight, the share of rows excluded for NULL or near-zero propensity, and the estimate under weight clipping at 20. State precisely what population your number describes.
Approach
- Audit before estimating. Split the rows into three buckets: usable (log_propensity strictly positive and recorded), NULL propensity, and positive but below a floor you choose and state. The NULL rows come from deterministic top-k serving, where no reweighting identifies the counterfactual, so they are not a data gap to impute; they are outside what this method can answer.
- Compute w = target_propensity / log_propensity on the usable rows. IPS is the mean of w * reward. It is unbiased under positivity and no unobserved confounding, and it has the variance problem that makes the rest of this exercise necessary.
- Compute SNIPS as sum(w * reward) / sum(w). It carries a small bias that vanishes with sample size, it is bounded inside the reward range so it cannot return an engagement rate above 1, and it is usually the number you would report.
- Report the effective sample size (sum w)^2 / sum(w^2) next to n. It is the honest denominator: 400,000 rows with ESS 3,100 is a 3,100-row estimate, and quoting the raw n next to a confidence interval derived from these weights is the way this analysis misleads people.
- Clip weights at the stated threshold, recompute, and describe the trade in the right direction: clipping caps variance and introduces downward bias wherever the target policy wants to act in regions the logging policy rarely visited, which is exactly where the candidate ranker differs most.
- State the estimand explicitly. After excluding the NULL and sub-floor rows, the number describes the sub-population of impressions where the logging policy explored, which is not the surface as a whole, and the decision that follows is whether to run an online test or add randomised exploration slots.
Worked solution 40 min
- Bucket the rows: usable = logged.log_propensity.notna() & (logged.log_propensity >= floor); report counts and impression share for usable, null, and below-floor
- u = logged[usable]; w = u.target_propensity / u.log_propensity
- ips = (w * u.reward).mean(); snips = (w * u.reward).sum() / w.sum()
- ess = w.sum()2 / (w2).sum(); p99 = w.quantile(0.99)
- w_clip = w.clip(upper=20); ips_clip = (w_clip * u.reward).mean(); snips_clip = (w_clip * u.reward).sum() / w_clip.sum()
- Write the estimand sentence naming the excluded share and the surface it no longer covers
Follow-up
- Sixty percent of rows have NULL log_propensity. What do you change about the serving system to make this analysis possible next quarter, and what does it cost?
- Add a doubly-robust estimator on top of this. What does the reward model buy you, and what happens when it is wrong?
- The offline estimate says plus 4 percent and the online test comes back flat. Give two mechanisms that produce exactly that pattern.
Write a SQL query using window functions to calculate rolling 7-day ac…
Write a SQL query using window functions to calculate rolling 7-day active user retention cohorts.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Optimize a slow-running query that aggregates daily messaging volumes …
Optimize a slow-running query that aggregates daily messaging volumes across multiple partitioned tables.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Split daily actives into new, retained, reactivated and resurrected
From daily_active(member_id, activity_date), one row per member per member-local day with a qualified session, classify every member active on day d into exactly one class: new (first-ever activity on d), retained (also active on d-1), reactivated (inactive on d-1 but active somewhere in d-28 to d-2), resurrected (inactive across all of d-28 to d-1 but active before d-28). Return day, class and member count; the four classes must sum to that day's active count. Report churn (active on d-1, absent on d) as a separate series, not a fifth class.
Approach
- Derive one value per row with LAG(activity_date) OVER (PARTITION BY member_id ORDER BY activity_date). Every class in the definition is a statement about the previous active day, so this single window function does all four.
- Map the definitions onto the gap: prev_date IS NULL means new; gap = 1 means retained; prev_date in [d-28, d-2], that is a gap of 2 through 28, means reactivated; prev_date earlier than d-28, a gap above 28, means resurrected. That equivalence is worth stating, because it is what makes the classes exhaustive rather than three predicates and an else.
- Write the classification as one CASE with branches in that fixed order, so exclusivity holds by construction. Four independent EXISTS subqueries look clearer and are the usual source of a member landing in two classes.
- Compute churn in a separate pass: day d-1's actives anti-joined to day d, indexed to the day the member went absent. Churn is not a class of actives on day d, since a churned member has no row on day d at all, and adding it would break the sum.
- Assert the check before showing anyone the result: for each day, the four class counts must sum to COUNT(DISTINCT member_id) in daily_active for that day. That check is the whole reason this decomposition is believed.
Worked solution 30 min
- WITH a AS (SELECT member_id, activity_date, LAG(activity_date) OVER (PARTITION BY member_id ORDER BY activity_date) AS prev_date FROM daily_active).
- Classify: CASE WHEN prev_date IS NULL THEN 'new' WHEN activity_date - prev_date = 1 THEN 'retained' WHEN activity_date - prev_date <= 28 THEN 'reactivated' ELSE 'resurrected' END.
- GROUP BY activity_date, class and count members.
- Churn: SELECT d.activity_date + 1 AS churn_date, COUNT(*) FROM daily_active d LEFT JOIN daily_active n ON n.member_id = d.member_id AND n.activity_date = d.activity_date + 1 WHERE n.member_id IS NULL GROUP BY 1.
- Run the sum check per day and only then join the class counts to the churn series for reporting.
Follow-up
- DAU is flat but the retained share fell four points while reactivated rose four. What happened, and what is the next query you run?
- How do these classes change if the day boundary is each member's local midnight rather than UTC midnight?
- Produce the same decomposition weekly instead of daily. Which class definition stops making sense?
Given a scenario with heavily skewed engagement metrics, which statist…
Given a scenario with heavily skewed engagement metrics, which statistical tests or transformations would you apply?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
Walk through how you would diagnose a sudden 15 percent drop in daily …
Walk through how you would diagnose a sudden 15 percent drop in daily active users for Snapchat Stories.
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How would you evaluate whether a new feature on the discovery tab cann…
How would you evaluate whether a new feature on the discovery tab cannibalizes existing engagement?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design a funnel dashboard and establish core metrics for…
How would you design a funnel dashboard and establish core metrics for a newly launched messaging feature?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
What experimentation pitfalls would you look out for if variant contam…
What experimentation pitfalls would you look out for if variant contamination occurs in a multi-region rollout?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
- State the primary metric and the minimum effect worth shipping, then size the test.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
How would you design a robust Bayesian A/B testing framework compared …
How would you design a robust Bayesian A/B testing framework compared to a frequentist approach?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Define a qualified session you can defend twice
The north-star counts weekly members with at least one qualified session. You own the definition. Today it reads: a session with 30 or more foreground seconds and at least one non-negative engagement or authored item, derived from fct_feed_impression.session_id, fct_feed_impression.dwell_ms, fct_engagement_event.action_type with is_negative_feedback = FALSE, and fct_content_item.created_at_utc. Write the definition precisely, including how a session closes and how undone actions count. Then defend it against two objections: that the 30-second floor excludes a real use, and that the engagement requirement is satisfiable by a mis-tap. Deliverable: the definition plus both rebuttals.
Approach
- Write the session boundary rule explicitly (inactivity gap, foreground only, how a background-to-foreground return is treated), because every rate in the tree is denominated in sessions and an unstated boundary makes the count a function of an SDK timeout nobody reviewed.
- Enumerate the qualifying acts from fct_engagement_event and the authored path from fct_content_item, and rule that an action with undone_at_utc within 60 seconds of occurred_at_utc does not qualify. That closes the mis-tap objection with data already logged rather than a new instrument.
- Answer the exclusion objection with evidence, not judgement: measure what fraction of sessions fall between 10 and 30 foreground seconds and what those members do the following week. Either accept the exclusion citing that number, or add a second qualifying path such as a completed deep-linked read, rather than lowering the floor globally.
- State what the definition deliberately is not: not time spent and not any impression, so a stickier or longer feed cannot move it without a member acting.
- Version the definition and pin it to a UI period, since moving the hide or like control changes the numerator with no change in member behaviour.
Worked solution 20 min
- Write the close rule: 30 minutes of foreground inactivity closes a session, a new foreground entry opens one, and a session is attributed to the viewer-local date of its first event using dim_member.tz_offset_minutes.
- Write the qualifying-act set: any fct_engagement_event row in the session with is_negative_feedback = FALSE and (undone_at_utc IS NULL OR undone_at_utc > occurred_at_utc + 60 seconds), or any fct_content_item authored by that member inside the session window.
- Run the sensitivity: recompute the weekly qualified-member count at foreground floors of 10, 30 and 60 seconds and record both the level and the week-over-week change under each.
- Write the two rebuttals, one paragraph each, each citing the specific number from step 3 that supports it.
Follow-up
- A platform team ships background prefetch that keeps the app foregrounded longer. Which part of your definition moves, and should it?
- You are asked to make this comparable across iOS, Android and web. What breaks first, and do you fix the definition or report the platforms separately?
Impressions dropped for three days, engagement events did not
Home_feed impressions on one client platform fell 22 percent for three days and then recovered. Engagement per 1,000 impressions on that platform rose 26 percent across the same three days and then returned to baseline. Raw engagement event counts were flat throughout. fct_engagement_event.impression_id is a nullable foreign key into fct_feed_impression, and warehouse foreign keys are declarative rather than enforced. Establish whether ranking improved or impression rows were lost, quantify the loss, and state the precondition your estimate depends on.
Approach
- Run the referential integrity test first, because it is decisive and cheap: count fct_engagement_event rows on those dates with impression_id IS NOT NULL that have no matching row in fct_feed_impression. Under normal operation this is near zero. Under impression loss it jumps, and each orphan is a directly observed missing impression.
- Check the reciprocal arithmetic. If the numerator is intact and the denominator loses a fraction f, the rate rises by 1/(1-f) - 1. An observed 26 percent rise implies f of about 0.206, which is consistent with the observed 22 percent drop within noise. A genuine ranking improvement has no reason to produce that near-exact reciprocal.
- Estimate total loss from the orphans: missing impressions is approximately orphan engagements divided by the baseline engagement-per-impression rate measured among matched rows on healthy days. State the precondition explicitly, that loss is independent of whether the impression was engaged with. If loss were correlated with engagement, this estimator is biased and the orphan count becomes a lower bound only.
- Characterise the loss as random or structured. Compare the distribution of impressions per session on the affected days against baseline: proportional loss shifts the whole distribution, while a failure that truncates a batch produces a distinctive spike of short sessions. Also check rows per hour to find the start and end of the window.
- Rule out a demand explanation. Sessions per member and engagement events per session should be flat if only logging broke. If sessions also fell, part of the impression drop is real and the two effects must be separated before quoting a loss figure.
- Recommend the operational response: mask these three dates in every impression-denominated rate, re-state them only if the upstream can replay, and note that any experiment reading on those dates is compromised for rate metrics but not for member counts.
Follow-up
- How do you decide whether to backfill the partitions or permanently mask the dates, and what does each choice cost downstream?
- An experiment was reading during those three days. Which of its metrics are still usable and which are not?
- What monitor would have caught this within an hour, and what is its false-positive cost?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.
Tell me about a time you had to deliver an unpopular data-driven findi…
Tell me about a time you had to deliver an unpopular data-driven finding to a resistant cross-functional stakeholder.
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Quantify the outcome, including what you would not claim credit for.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Walk through an analysis you shipped that was wrong
Prepare a five-minute account of an analysis you delivered that a decision was made on and that later turned out to be wrong. Cover the conclusion you gave, the mechanism of the error stated precisely enough that a peer could reproduce it, how and by whom it was found, how long it stood, what the wrong decision cost, and the specific check now in your workflow. Expect the interviewer to ask which of your current results is most likely wrong for the same reason.
Approach
- Lead with the decision and the error in two sentences, then go back for detail. An account that opens with context loses the interviewer before the mechanism arrives.
- State the mechanism as a data fact, not a mood. 'A repeated delivery of the same item to the same viewer counted as two impressions in the denominator, so the engagement rate was deflated for high-redelivery surfaces' is assessable. 'The data was messy' is not, and it reads as not having understood the bug.
- Say who found it without softening. If a reviewer or a downstream team found it, say what in your process let it through: a rate rolled up by averaging sub-period rates, a filter applied after treatment, a class decomposition never checked against its total.
- Quantify the cost in the currency of the decision rather than in revenue you cannot support: a launch held for six weeks, a team quarter spent on the wrong lever, a metric definition that ten later decisions inherited. Then say how much of the original conclusion survived the correction, because often the direction held and only the magnitude broke.
- Close on a check that either runs or does not: printing row counts at each grain before aggregating, asserting that the four DAU classes sum to DAU, reporting numerator and denominator beside every ratio. Give one instance of that check firing since, which is what separates a process change from an intention.
Follow-up
- What did you tell the people who had already acted on the wrong number, and when?
- Which result you currently stand behind is most likely wrong for the same reason?
- Why was that check not already in the work? What made it feel unnecessary at the time?
Explain a prevalence interval to a non-technical executive
A weekly impression-weighted violating-content prevalence estimate came in at 0.42 percent, 95 percent interval 0.28 to 0.61, against 0.51 percent (0.35 to 0.72) the week before. The audit sample is 4,000 served impressions drawn with unequal, recorded selection probabilities across risk strata, labelled by humans against written policy. An executive asks whether the number went down and wants one figure for a board slide. In five minutes: answer the question, say what goes on the slide, and state what you would need to give a sharper answer next quarter.
Approach
- Answer the question in one sentence before explaining anything: the point estimate is lower, the intervals overlap across most of their range, and the week-over-week change is not distinguishable from zero.
- Show why with one arithmetic step rather than vocabulary. At n = 4,000 and p near 0.004 the simple-random-sampling standard error is sqrt(p(1-p)/n), about 0.10 percentage points, so an SRS interval would run roughly plus or minus 0.20 points and a 0.09 point move sits well inside it. Two facts about the reported interval belong in your head rather than on the slide. Its asymmetry comes from the construction, not from the weights: Wilson, Clopper-Pearson and logit intervals are built on a bounded scale, so near p = 0 the upper limit sits further from the point estimate than the lower one. The 1/p_i weights act on width only, through a design effect that multiplies the variance. Here the reported width of 0.33 points implies a standard error near 0.085 (0.33 divided by 3.92), so the design effect is about 0.7, which is what oversampling high-risk strata buys when selection probability correlates with the outcome. Uninformative weights would instead give a design effect of 1 + CV squared of the weights, above 1, and an interval wider than the SRS one rather than narrower.
- Replace the bare point estimate with a number that is stable at board cadence: the trailing four-week pooled estimate, formed by re-summing the weighted numerator and the weighted denominator across weeks. Averaging the four weekly rates gives a different and wrong number when weekly sample sizes differ.
- Price the precision the executive is implicitly asking for. Halving the interval width needs roughly four times the labelled sample, so 16,000 labels a week to go from a half-width near 0.17 points to one near 0.085. The cheaper lever is allocation rather than volume: the design already uses unequal, recorded, strictly positive selection probabilities and is already running a design effect near 0.7, so re-fitting the strata on current classifier scores and moving more of the 4,000 into the strata carrying the violating mass pushes that number down further without a fourfold labelling bill.
- State plainly what this number is not, because the executive will meet substitutes. Report volume and enforcement volume are member and operations behaviours; they can fall while prevalence rises if the ranker gets better at matching violating content to receptive audiences.
Follow-up
- The executive wants a weekly trend line on the slide anyway. What do you draw, and what do you label the band?
- How long would it take to detect a 20 percent reduction in prevalence at the current sample size?
- Why not score every impression with the classifier instead of paying for human labels?
- 01
Tell me about a time you had to deliver an unpopular data-driven finding to a resistant cross-functional stakeholder.
- 02
Prepare a five-minute account of an analysis you delivered that a decision was made on and that later turned out to be wrong. Cover the conclusion you gave, the mechanism of the error stated precisely enough that a peer could reproduce it, how and by whom it was found, how long it stood, what the wrong decision cost, and the specific check now in your workflow. Expect the interviewer to ask which of your current results is most likely wrong for the same reason.
- 03
A weekly impression-weighted violating-content prevalence estimate came in at 0.42 percent, 95 percent interval 0.28 to 0.61, against 0.51 percent (0.35 to 0.72) the week before. The audit sample is 4,000 served impressions drawn with unequal, recorded selection probabilities across risk strata, labelled by humans against written policy. An executive asks whether the number went down and wants one figure for a board slide. In five minutes: answer the question, say what goes on the slide, and state what you would need to give a sharper answer next quarter.
Is this an official Snap interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Snap. Rounds and questions reflect what candidates have reported, not a process Snap has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the overall difficulty level of the interview process?
The interview loop is rigorous and comparable to top-tier technology companies. Expect thorough technical screens followed by a comprehensive onsite loop testing your coding, statistics, and product sense under pressure.
PracHub interview research ↗How should I prepare for the product sense and metrics rounds?
Focus on structuring ambiguous problems by first clarifying goals, defining user cohorts, establishing primary and guardrail metrics, and outlining a step-by-step diagnostic framework for metric fluctuations.
PracHub interview research ↗Are there LeetCode-style algorithms tested in the coding rounds?
Heavy data structures and algorithms questions are generally rare for this role. Instead, expect pragmatic Python coding challenges focused on data manipulation, numerical computation, and basic logic.
PracHub interview research ↗What is the workplace flexibility policy at the company?
The organization follows a default together approach, expecting team members to work from an office multiple days per week to foster dynamic collaboration and build culture faster.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22