Snap · Data Scientist
Updated · 2026-09-22

Snap Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Snap, you sit at the intersection of quantitative rigor and high-velocity product innovation. You drive the analytical engine behind core ecosystem products like Snapchat, Lens Studio, and Spectacles, translating massive, complex datasets into clear, actionable business strategies. Your work directly shapes how millions of people globally communicate, express themselves, and interact with augmented reality.

When randomisation is off the table, the skill being checked is naming an identification strategy together with the assumption it rests on: parallel trends for difference-in-differences, relevance and exclusion for an instrument, overlap and conditional ignorability for matching. Say the assumption out loud and say how you would try to break it.

Snap candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Design feed experiments that survive social-graph interferenceDecompose DAU into new, retained, reactivated, resurrectedAggregate impression logs without double-counting repeated deliveries

35 min read

Practice 17 Data Scientist prompts
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Snap, you sit at the intersection of quantitative rigor and high-velocity product innovation. You drive the analytical engine behind core ecosystem products like Snapchat, Lens Studio, and Spectacles, translating massive, complex datasets into clear, actionable business strategies. Your work directly shapes how millions of people globally communicate, express themselves, and interact with augmented reality.

This role requires you to act as both a truth-seeker and a strategic advisor. Whether you are optimizing monetization funnels, scaling growth loops, or evaluating new messaging features, your insights steer roadmap prioritization across engineering, product management, and design. You will tackle ambiguous, high-impact problems where defining the right metric is just as important as building the predictive model behind it.

Expect an environment characterized by rapid experimentation and deep technical ownership. You will not just pull numbers; you will design rigorous A/B tests, diagnose unexpected metric fluctuations, and deploy scalable statistical solutions. If you thrive on ambiguity, possess strong product intuition, and want your analysis to directly influence a globally recognized product, this role offers an ideal platform.

01

Recruiter Conversation

reported

A screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.

What to demonstrate

  • Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
  • Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
  • Whether timeline, location and compensation expectations make the rest of the loop worth scheduling

How to prepare

  • Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
  • Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
  • Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
PracHub interview research ↗
02

Technical Screen

reported

A handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.

What to demonstrate

  • Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
  • Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
  • Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.

How to prepare

  • Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
  • Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
  • If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
PracHub interview research ↗
03

Virtual Onsite Loop

reported

Where a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.

What to demonstrate

  • Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
  • Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
  • Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
  • Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered

How to prepare

  • Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
  • Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
  • Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Crediting notifications with the sessions that follow them

Members who open a push notification were already more likely to open the app, so attributing their session to the notification measures intent rather than causation and makes almost any send look profitable. The predictable result is a push-volume increase that shows a large modelled gain and a small real one, paid for later in opt-outs and uninstalls that no single experiment window captures. The only defensible estimate compares a stable send arm against a holdout arm assigned at the decision point, with the held-out decisions logged and suppressed at delivery, over a window long enough to include the opt-out response. Frequency effects are non-linear, so a per-notification incremental rate estimated at one volume does not extrapolate to a higher one.

02

Using report volume as a measurement of how much violating content exists

Reporting is a member behaviour, not an observation of the content. Report counts rise when the report control is made easier to reach, when a coordinated campaign targets an account, and when the audience shifts toward people who object; they fall when violating content is shown mainly to members who agree with it. A ranker that gets better at matching bad content to receptive audiences will drive reports down and harm up at the same time. Prevalence must come from a random sample of served impressions with recorded selection probabilities, labelled by humans against the written policy, and reported with an interval. Reports are useful as a detection signal and as a demand-side complaint rate, not as a denominator-anchored measure of harm.

03

Answering a product-sense question with a list of features

Answer with a decision and the measurement that would settle it: the hypothesis, the primary metric, the guardrails, and the result that would make you not ship. A feature brainstorm cannot be wrong, which is exactly why it earns no points.

04

Reading an observational correlation as a causal effect

Name the confounder you are most worried about and the design that would remove it: an experiment, a difference-in-differences with a checked pre-period trend, an instrument, or a regression discontinuity. When none is available, state which direction the bias likely runs and bound the claim accordingly.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

How do you test for heterogeneous treatment effects in an offline data…

medium
statistics and probability

How do you test for heterogeneous treatment effects in an offline data sample?

Approach
  1. Sanity-check the answer against a simple bound or a simulated case.
  2. Say what the estimate is of, and over what population it generalises.
  3. Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • How would you explain this result to someone who does not know statistics?

Explain the core assumptions behind linear regression and how you iden…

medium
machine learning and modelling

Explain the core assumptions behind linear regression and how you identify violations in observational data.

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Say how the offline result would be validated online before it is trusted.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

Off-policy estimate with a positivity and weight audit

hardWorked solution
off-policy evaluationpropensity weightspositivity

logged has impression_id, rank_position, reward (1 if the viewer engaged), log_propensity (float, NULL where serving was deterministic top-k) and target_propensity (the candidate policy's probability of placing the same item in the same slot, precomputed). Estimate the candidate policy's engagement rate per impression with inverse propensity scoring and with self-normalised IPS. Report the effective sample size as the square of the sum of weights over the sum of squared weights, the 99th percentile weight, the share of rows excluded for NULL or near-zero propensity, and the estimate under weight clipping at 20. State precisely what population your number describes.

Approach
  1. Audit before estimating. Split the rows into three buckets: usable (log_propensity strictly positive and recorded), NULL propensity, and positive but below a floor you choose and state. The NULL rows come from deterministic top-k serving, where no reweighting identifies the counterfactual, so they are not a data gap to impute; they are outside what this method can answer.
  2. Compute w = target_propensity / log_propensity on the usable rows. IPS is the mean of w * reward. It is unbiased under positivity and no unobserved confounding, and it has the variance problem that makes the rest of this exercise necessary.
  3. Compute SNIPS as sum(w * reward) / sum(w). It carries a small bias that vanishes with sample size, it is bounded inside the reward range so it cannot return an engagement rate above 1, and it is usually the number you would report.
  4. Report the effective sample size (sum w)^2 / sum(w^2) next to n. It is the honest denominator: 400,000 rows with ESS 3,100 is a 3,100-row estimate, and quoting the raw n next to a confidence interval derived from these weights is the way this analysis misleads people.
  5. Clip weights at the stated threshold, recompute, and describe the trade in the right direction: clipping caps variance and introduces downward bias wherever the target policy wants to act in regions the logging policy rarely visited, which is exactly where the candidate ranker differs most.
  6. State the estimand explicitly. After excluding the NULL and sub-floor rows, the number describes the sub-population of impressions where the logging policy explored, which is not the surface as a whole, and the decision that follows is whether to run an online test or add randomised exploration slots.
Worked solution 40 min
  1. Bucket the rows: usable = logged.log_propensity.notna() & (logged.log_propensity >= floor); report counts and impression share for usable, null, and below-floor
  2. u = logged[usable]; w = u.target_propensity / u.log_propensity
  3. ips = (w * u.reward).mean(); snips = (w * u.reward).sum() / w.sum()
  4. ess = w.sum()2 / (w2).sum(); p99 = w.quantile(0.99)
  5. w_clip = w.clip(upper=20); ips_clip = (w_clip * u.reward).mean(); snips_clip = (w_clip * u.reward).sum() / w_clip.sum()
  6. Write the estimand sentence naming the excluded share and the surface it no longer covers
EXPECTED RESULTFive numbers with one sentence of scope: IPS, SNIPS, ESS against n, the 99th percentile weight, and the clipped SNIPS, plus the excluded share split into NULL and below-floor, and an explicit statement that the estimate applies to the explored sub-population only.
Follow-up
  • Sixty percent of rows have NULL log_propensity. What do you change about the serving system to make this analysis possible next quarter, and what does it cost?
  • Add a doubly-robust estimator on top of this. What does the reward model buy you, and what happens when it is wrong?
  • The offline estimate says plus 4 percent and the online test comes back flat. Give two mechanisms that produce exactly that pattern.

For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Design one test end to end on paper
  • Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
  • Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
  • State in advance what you will do if the primary metric is flat while a secondary metric is significant.

Deliverable: A one-page test design with a decision rule written before launch.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Power arithmetic until it is automatic
  • Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
  • Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
  • Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.

Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Variance and the unit-of-analysis problem
  • Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
  • Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
  • Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.

Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Validity threats you can actually test for
  • Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
  • Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
  • Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.

Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05When randomization is not available
  • Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
  • Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
  • List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.

Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.

Practice prompt ↗Practice prompt ↗
06The readout query
  • Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
  • Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
  • Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.

Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.

Practice prompt ↗Practice prompt ↗
07Present it to someone who will not read the appendix
  • Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
  • Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
  • Rewrite your opening line so the recommendation lands before any methodology.

Deliverable: A one-page readout whose first line is the recommendation.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.

Tell me about a time you had to deliver an unpopular data-driven findi…

medium
behavioural and stakeholder questions

Tell me about a time you had to deliver an unpopular data-driven finding to a resistant cross-functional stakeholder.

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Quantify the outcome, including what you would not claim credit for.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

Walk through an analysis you shipped that was wrong

medium
postmortemgrain errorscredibility

Prepare a five-minute account of an analysis you delivered that a decision was made on and that later turned out to be wrong. Cover the conclusion you gave, the mechanism of the error stated precisely enough that a peer could reproduce it, how and by whom it was found, how long it stood, what the wrong decision cost, and the specific check now in your workflow. Expect the interviewer to ask which of your current results is most likely wrong for the same reason.

Approach
  1. Lead with the decision and the error in two sentences, then go back for detail. An account that opens with context loses the interviewer before the mechanism arrives.
  2. State the mechanism as a data fact, not a mood. 'A repeated delivery of the same item to the same viewer counted as two impressions in the denominator, so the engagement rate was deflated for high-redelivery surfaces' is assessable. 'The data was messy' is not, and it reads as not having understood the bug.
  3. Say who found it without softening. If a reviewer or a downstream team found it, say what in your process let it through: a rate rolled up by averaging sub-period rates, a filter applied after treatment, a class decomposition never checked against its total.
  4. Quantify the cost in the currency of the decision rather than in revenue you cannot support: a launch held for six weeks, a team quarter spent on the wrong lever, a metric definition that ten later decisions inherited. Then say how much of the original conclusion survived the correction, because often the direction held and only the magnitude broke.
  5. Close on a check that either runs or does not: printing row counts at each grain before aggregating, asserting that the four DAU classes sum to DAU, reporting numerator and denominator beside every ratio. Give one instance of that check firing since, which is what separates a process change from an intention.
Follow-up
  • What did you tell the people who had already acted on the wrong number, and when?
  • Which result you currently stand behind is most likely wrong for the same reason?
  • Why was that check not already in the work? What made it feel unnecessary at the time?

Explain a prevalence interval to a non-technical executive

easy
uncertaintyprevalenceexecutive communication

A weekly impression-weighted violating-content prevalence estimate came in at 0.42 percent, 95 percent interval 0.28 to 0.61, against 0.51 percent (0.35 to 0.72) the week before. The audit sample is 4,000 served impressions drawn with unequal, recorded selection probabilities across risk strata, labelled by humans against written policy. An executive asks whether the number went down and wants one figure for a board slide. In five minutes: answer the question, say what goes on the slide, and state what you would need to give a sharper answer next quarter.

Approach
  1. Answer the question in one sentence before explaining anything: the point estimate is lower, the intervals overlap across most of their range, and the week-over-week change is not distinguishable from zero.
  2. Show why with one arithmetic step rather than vocabulary. At n = 4,000 and p near 0.004 the simple-random-sampling standard error is sqrt(p(1-p)/n), about 0.10 percentage points, so an SRS interval would run roughly plus or minus 0.20 points and a 0.09 point move sits well inside it. Two facts about the reported interval belong in your head rather than on the slide. Its asymmetry comes from the construction, not from the weights: Wilson, Clopper-Pearson and logit intervals are built on a bounded scale, so near p = 0 the upper limit sits further from the point estimate than the lower one. The 1/p_i weights act on width only, through a design effect that multiplies the variance. Here the reported width of 0.33 points implies a standard error near 0.085 (0.33 divided by 3.92), so the design effect is about 0.7, which is what oversampling high-risk strata buys when selection probability correlates with the outcome. Uninformative weights would instead give a design effect of 1 + CV squared of the weights, above 1, and an interval wider than the SRS one rather than narrower.
  3. Replace the bare point estimate with a number that is stable at board cadence: the trailing four-week pooled estimate, formed by re-summing the weighted numerator and the weighted denominator across weeks. Averaging the four weekly rates gives a different and wrong number when weekly sample sizes differ.
  4. Price the precision the executive is implicitly asking for. Halving the interval width needs roughly four times the labelled sample, so 16,000 labels a week to go from a half-width near 0.17 points to one near 0.085. The cheaper lever is allocation rather than volume: the design already uses unequal, recorded, strictly positive selection probabilities and is already running a design effect near 0.7, so re-fitting the strata on current classifier scores and moving more of the 4,000 into the strata carrying the violating mass pushes that number down further without a fourfold labelling bill.
  5. State plainly what this number is not, because the executive will meet substitutes. Report volume and enforcement volume are member and operations behaviours; they can fall while prevalence rises if the ranker gets better at matching violating content to receptive audiences.
Follow-up
  • The executive wants a weekly trend line on the slide anyway. What do you draw, and what do you label the band?
  • How long would it take to detect a 20 percent reduction in prevalence at the current sample size?
  • Why not score every impression with the classifier instead of paying for human labels?
  • 01

    Tell me about a time you had to deliver an unpopular data-driven finding to a resistant cross-functional stakeholder.

  • 02

    Prepare a five-minute account of an analysis you delivered that a decision was made on and that later turned out to be wrong. Cover the conclusion you gave, the mechanism of the error stated precisely enough that a peer could reproduce it, how and by whom it was found, how long it stood, what the wrong decision cost, and the specific check now in your workflow. Expect the interviewer to ask which of your current results is most likely wrong for the same reason.

  • 03

    A weekly impression-weighted violating-content prevalence estimate came in at 0.42 percent, 95 percent interval 0.28 to 0.61, against 0.51 percent (0.35 to 0.72) the week before. The audit sample is 4,000 served impressions drawn with unequal, recorded selection probabilities across risk strata, labelled by humans against written policy. An executive asks whether the number went down and wants one figure for a board slide. In five minutes: answer the question, say what goes on the slide, and state what you would need to give a sharper answer next quarter.

PracHub interview preparation framework ↗
Is this an official Snap interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Snap. Rounds and questions reflect what candidates have reported, not a process Snap has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
What is the overall difficulty level of the interview process?

The interview loop is rigorous and comparable to top-tier technology companies. Expect thorough technical screens followed by a comprehensive onsite loop testing your coding, statistics, and product sense under pressure.

PracHub interview research ↗
How should I prepare for the product sense and metrics rounds?

Focus on structuring ambiguous problems by first clarifying goals, defining user cohorts, establishing primary and guardrail metrics, and outlining a step-by-step diagnostic framework for metric fluctuations.

PracHub interview research ↗
Are there LeetCode-style algorithms tested in the coding rounds?

Heavy data structures and algorithms questions are generally rare for this role. Instead, expect pragmatic Python coding challenges focused on data manipulation, numerical computation, and basic logic.

PracHub interview research ↗
What is the workplace flexibility policy at the company?

The organization follows a default together approach, expecting team members to work from an office multiple days per week to foster dynamic collaboration and build culture faster.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.