As a Data Scientist at Nextdoor, you serve as a core mathematical decision scientist within a semi-embedded product development organization. You collaborate directly with product, engineering, design, and operations stakeholders to shape how millions of neighbors connect across local communities worldwide. Your work spans critical areas such as search relevance, recommendation systems, trust and safety, and marketplace dynamics, directly influencing how users discover local information, businesses, and each other.
The role combines deep analytical rigor with entrepreneurial execution, functioning as an analyst and a builder simultaneously. You will dive into large-scale, complex datasets to uncover behavioral insights, design robust experimentation frameworks, and prototype machine learning solutions that operate in a modern AI-first environment. Whether you are investigating metric shifts, optimizing ranking algorithms, or evaluating platform safety policies, your insights translate directly into product strategy and tangible user impact.
Expect a fast-paced and collaborative culture where technical excellence meets a clear social mission. operates with a lean and hungry data science community, giving individual contributors high visibility and substantial ownership over major product surfaces. Success in this role requires a balance of sophisticated technical skills, strong product intuition, and the communication prowess to translate complex quantitative findings into clear strategic narratives.
Initial Screening
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Assessments
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Behavioral Interviews
reportedMost of the weight in this round sits on the disagreement questions. Data work routinely produces an answer someone senior did not want, and the interviewer is trying to learn what you do in that hour. Both failure modes are common: folding as soon as a director pushes back, and treating the pushback as ignorance to be corrected with a better chart. A strong answer usually contains a specific thing the other person knew that you did not, and describes how you found out whether it changed the conclusion.
What to demonstrate
- Whether you can state the other side's argument accurately before you explain why you disagreed
- What you treated as evidence during the disagreement, such as a rerun under their assumption or a holdout check, rather than persuasion technique
- Whether you distinguish being overruled from being wrong, and can give an example of each
How to prepare
- Write out one disagreement where you turned out to be wrong, and say what in the data misled you. Candidates prepare the story where they were right, and the follow-up asks for the other one.
- For your main disagreement story, be ready to say what result would have made you drop your position. If no such result exists, you were not arguing from the data.
- Practise stating the opposing position out loud in one sentence the stakeholder would accept, then continue the story.
Case Study Discussions
reportedA case round is decided by whether you leave the interviewer with a recommendation, not by how much analysis you narrate on the way there. The prompt is open on purpose, so the first job is to convert it into a decision someone could act on: ask what would be done differently depending on the answer. From there name the quantity that would settle it, state the assumptions you need, and commit. Candidates who cover more ground than anyone expected and still end on "it depends" score below candidates who scoped narrowly and said what they would do.
What to demonstrate
- Whether the version of the question you choose to answer is genuinely narrower than the prompt and still worth answering
- Whether the recommendation arrives as an action with a number attached, rather than as a summary of what you looked at
- Whether assumptions are stated at the moment you rely on them, instead of collected into a disclaimer at the end
- Whether you notice when a branch you are exploring would not change the decision either way
How to prepare
- Take six open prompts and write only the scoping move for each: the one-sentence question you would actually answer and the decision it feeds. Give yourself three minutes per prompt and stop there.
- Put a five-minute warning into every practice case and force a closing statement that names the action, the result that would justify it, and the result that would reverse it.
- Record one case and count how long you talked before naming a measurable quantity. Past roughly five minutes, what you are calling scoping is narration.
Final Interviews
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
1 candidate reports. Individual accounts describe a particular role and hiring cycle.
Nextdoor Data Scientist Interview Experience — A Surprise Case Study, Then the Role Was Paused
View report detailsPracHub editorial advice for the preparation topics above.
Crediting notifications with the sessions that follow them
Members who open a push notification were already more likely to open the app, so attributing their session to the notification measures intent rather than causation and makes almost any send look profitable. The predictable result is a push-volume increase that shows a large modelled gain and a small real one, paid for later in opt-outs and uninstalls that no single experiment window captures. The only defensible estimate compares a stable send arm against a holdout arm assigned at the decision point, with the held-out decisions logged and suppressed at delivery, over a window long enough to include the opt-out response. Frequency effects are non-linear, so a per-notification incremental rate estimated at one volume does not extrapolate to a higher one.
Averaging per-member engagement when the distribution is heavy-tailed
Time spent, impressions and authored actions per member are extremely right-skewed, so the sample mean has high variance and a handful of members can move a reported lift past significance on their own. Two specific errors follow. Averaging per-member rates (mean of ratios) answers a different question from total engagements over total members (ratio of sums), and the two diverge whenever the heavy users behave differently. Capping or winsorising reduces variance but changes the estimand, so the cap must be pre-registered and reported, not chosen after seeing the result. Ratio metrics with the member as randomisation unit also need delta-method or bootstrap variance, since numerator and denominator are correlated.
Dropping rows with missing values without naming the mechanism
Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.
Reporting a mean for a heavy-tailed metric without saying what it hides
For spend, session length or items per order, a small fraction of units carries most of the total, so the mean has a wide standard error and one account can move it. Fix the handling before you see the result: cap or winsorise at a pre-declared percentile, and report the median or the share above a threshold next to the mean. Capping changes the estimand, so say which question the capped number answers, and check how much of any difference comes from the top 0.1 percent of units.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
There is something strange about this dataset. Can you find it?
There is something strange about this dataset. Can you find it?
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Sanity-check the answer against a simple bound or a simulated case.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
Walk through the end-to-end design of a machine learning model to proa…
Walk through the end-to-end design of a machine learning model to proactively detect fraudulent posts or abusive behavior.
Approach
- Say how the offline result would be validated online before it is trusted.
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Sessionise an event stream with a thirty-minute inactivity gap
events has member_id (int64), occurred_at_utc (datetime64[ns], UTC) and action_kind, roughly 5 million rows in arbitrary order. Close a session after 30 minutes of inactivity for that member. Assign session_id using vectorised pandas only: no groupby.apply, no Python loop over members. Then compute per session the span in seconds between first and last event, the event count, and whether the session qualifies (span of at least 30 seconds and at least one non-negative engagement or authored item). Return the session table and the count of distinct members with at least one qualified session.
Approach
- Sort by (member_id, occurred_at_utc) once. Every later step assumes that order, so do it explicitly rather than relying on the input arriving sorted.
- Compute the gap as events.groupby('member_id').occurred_at_utc.diff(). The groupby is the whole point: a plain .diff() over the frame measures the gap between the last event of one member and the first of the next, which silently merges two members into one session at every boundary.
- new_session = gap.isna() | (gap > Timedelta('30min')). The isna arm opens the first session of each member. Decide and state whether a gap of exactly 30 minutes opens a new session; either convention is fine but it has to be written down because it moves the count.
- session_ordinal = new_session.groupby(events.member_id).cumsum(), then build a session key from (member_id, ordinal) with factorize so the id is a compact int rather than a string concat over 5 million rows.
- Aggregate once with a single groupby(session_key).agg: min and max timestamp, size, and a boolean any over the qualifying action mask. Derive span_seconds from the aggregated min and max, not row by row.
- Flag the known weakness out loud: span between first and last event is a lower bound on foreground time, and a one-event session gets span 0, so it can never qualify under a 30-second rule. That is a definitional choice about single-event sessions, not a bug.
Worked solution 30 min
- ev = events.sort_values(['member_id','occurred_at_utc'], kind='mergesort').reset_index(drop=True)
- gap = ev.groupby('member_id', sort=False).occurred_at_utc.diff(); new = gap.isna() | (gap > pd.Timedelta(minutes=30))
- ordinal = new.groupby(ev.member_id, sort=False).cumsum(); ev['session_key'] = pd.factorize(pd.MultiIndex.from_arrays([ev.member_id, ordinal]))[0]
- ev['is_qualifying_action'] = ev.action_kind.isin(QUALIFYING); s = ev.groupby('session_key').agg(member_id=('member_id','first'), t0=('occurred_at_utc','min'), t1=('occurred_at_utc','max'), n=('session_key','size'), any_action=('is_qualifying_action','any'))
- s['span_s'] = (s.t1 - s.t0).dt.total_seconds(); s['qualified'] = (s.span_s >= 30) & s.any_action
- s.loc[s.qualified, 'member_id'].nunique()
Follow-up
- Where does the 30-minute threshold come from, and how would you choose it from the data rather than inheriting it?
- The client emits a heartbeat every 10 seconds while the app is foregrounded. How does that change both the session boundaries and the span calculation?
- A member has two devices active at once. What does your session_id mean now, and does the qualified-session count still answer the question it is used for?
Write a query using SQL window functions to calculate rolling 7-day ac…
Write a query using SQL window functions to calculate rolling 7-day active user engagement across neighborhoods.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Clean and format unstructured text metadata for preliminary search rel…
Clean and format unstructured text metadata for preliminary search relevance evaluation.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Split daily actives into new, retained, reactivated and resurrected
From daily_active(member_id, activity_date), one row per member per member-local day with a qualified session, classify every member active on day d into exactly one class: new (first-ever activity on d), retained (also active on d-1), reactivated (inactive on d-1 but active somewhere in d-28 to d-2), resurrected (inactive across all of d-28 to d-1 but active before d-28). Return day, class and member count; the four classes must sum to that day's active count. Report churn (active on d-1, absent on d) as a separate series, not a fifth class.
Approach
- Derive one value per row with LAG(activity_date) OVER (PARTITION BY member_id ORDER BY activity_date). Every class in the definition is a statement about the previous active day, so this single window function does all four.
- Map the definitions onto the gap: prev_date IS NULL means new; gap = 1 means retained; prev_date in [d-28, d-2], that is a gap of 2 through 28, means reactivated; prev_date earlier than d-28, a gap above 28, means resurrected. That equivalence is worth stating, because it is what makes the classes exhaustive rather than three predicates and an else.
- Write the classification as one CASE with branches in that fixed order, so exclusivity holds by construction. Four independent EXISTS subqueries look clearer and are the usual source of a member landing in two classes.
- Compute churn in a separate pass: day d-1's actives anti-joined to day d, indexed to the day the member went absent. Churn is not a class of actives on day d, since a churned member has no row on day d at all, and adding it would break the sum.
- Assert the check before showing anyone the result: for each day, the four class counts must sum to COUNT(DISTINCT member_id) in daily_active for that day. That check is the whole reason this decomposition is believed.
Worked solution 30 min
- WITH a AS (SELECT member_id, activity_date, LAG(activity_date) OVER (PARTITION BY member_id ORDER BY activity_date) AS prev_date FROM daily_active).
- Classify: CASE WHEN prev_date IS NULL THEN 'new' WHEN activity_date - prev_date = 1 THEN 'retained' WHEN activity_date - prev_date <= 28 THEN 'reactivated' ELSE 'resurrected' END.
- GROUP BY activity_date, class and count members.
- Churn: SELECT d.activity_date + 1 AS churn_date, COUNT(*) FROM daily_active d LEFT JOIN daily_active n ON n.member_id = d.member_id AND n.activity_date = d.activity_date + 1 WHERE n.member_id IS NULL GROUP BY 1.
- Run the sum check per day and only then join the class counts to the churn series for reporting.
Follow-up
- DAU is flat but the retained share fell four points while reactivated rose four. What happened, and what is the next query you run?
- How do these classes change if the day boundary is each member's local midnight rather than UTC midnight?
- Produce the same decomposition weekly instead of daily. Which class definition stops making sense?
Define a comprehensive framework of core and guardrail metrics for a n…
Define a comprehensive framework of core and guardrail metrics for a newly launched local business search feature.
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
Questions involved product experience. Examples of a question could be…
Questions involved product experience. Examples of a question could be, how would you investigate a metric dropping X% amount…
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Preprocess and aggregate geospatial activity logs to identify regional…
Preprocess and aggregate geospatial activity logs to identify regional community growth trends.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Extract and aggregate session-level retention metrics from a large eve…
Extract and aggregate session-level retention metrics from a large event log table.
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you handle network effects and interference between treatmen…
How would you handle network effects and interference between treatment and control units in a local social network experiment?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
How do you account for multiple testing corrections when evaluating do…
How do you account for multiple testing corrections when evaluating dozens of secondary metrics in an experiment?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- What would you do if you could not randomise at all?
How you measure a product success
How you measure a product success
Approach
- Say what you would check first and why it is the highest-information step.
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Diagnose a sample ratio mismatch before reading the lift
The dashboard for a 50/50 home_feed test reports, over 9 days, 402,913 distinct viewer_member_id in control and 398,412 in treatment, read from fct_feed_impression.experiment_bucket. Impression rows differ between arms by 4 percent. The team is ready to read a +1.4 percent lift on the primary. Deliver the test statistic and your verdict on whether the lift is readable, the three cuts you would run first to localise the cause, and whether the impression-row gap is itself evidence of a problem.
Approach
- Test at the randomisation unit only. Chi-square goodness of fit on distinct assigned members against the designed 50/50 split, 1 degree of freedom, alerting at a strict threshold such as p < 0.001 because the test runs on every experiment every day and a 0.05 alarm fires constantly.
- Rule the impression gap out as evidence: impressions per member is an outcome the treatment is designed to move, so an arm difference there is expected and carries no information about assignment integrity.
- Localise by splitting the chi-square: by event_date to find the onset day, by client_platform and app_version to catch an arm-specific crash or a build that fails to log the bucket, and by tenure or registration date to catch members who enter the experiment through a different code path.
- Interrogate the trigger point. If the bucket is stamped at first impression rather than at assignment, the treatment can change who ever gets stamped, which produces mismatch and differential triggering at once; reconcile the impression-derived counts against the assignment service log.
- Check analysis-side filters applied after assignment: bot exclusion, impression dedupe, dropping suspended or deleted accounts. Any of these applied asymmetrically produces the same signature.
- Fix the cause and restart. Do not reweight the arms to the designed ratio; the missing members are not missing at random with respect to the outcome.
Worked solution 15 min
- Total assigned = 402,913 + 398,412 = 801,325, so each arm is expected to hold 400,662.5.
- Deviation is 2,250.5 per arm; chi-square = 2 x 2,250.5^2 / 400,662.5.
- Convert to a p-value on 1 degree of freedom and compare against the alert threshold.
- State the observed split as a percentage so the size of the problem is legible to non-statisticians.
- List the three cuts and the one reconciliation against the assignment log.
Follow-up
- The imbalance is confined to one app_version. Can you analyse the remaining versions and ship?
- What alert threshold do you set for SRM across hundreds of concurrent tests, and how do you keep the false alarm rate tolerable?
- An SRM appears only after day 6. What single hypothesis does that timing favour?
Impressions dropped for three days, engagement events did not
Home_feed impressions on one client platform fell 22 percent for three days and then recovered. Engagement per 1,000 impressions on that platform rose 26 percent across the same three days and then returned to baseline. Raw engagement event counts were flat throughout. fct_engagement_event.impression_id is a nullable foreign key into fct_feed_impression, and warehouse foreign keys are declarative rather than enforced. Establish whether ranking improved or impression rows were lost, quantify the loss, and state the precondition your estimate depends on.
Approach
- Run the referential integrity test first, because it is decisive and cheap: count fct_engagement_event rows on those dates with impression_id IS NOT NULL that have no matching row in fct_feed_impression. Under normal operation this is near zero. Under impression loss it jumps, and each orphan is a directly observed missing impression.
- Check the reciprocal arithmetic. If the numerator is intact and the denominator loses a fraction f, the rate rises by 1/(1-f) - 1. An observed 26 percent rise implies f of about 0.206, which is consistent with the observed 22 percent drop within noise. A genuine ranking improvement has no reason to produce that near-exact reciprocal.
- Estimate total loss from the orphans: missing impressions is approximately orphan engagements divided by the baseline engagement-per-impression rate measured among matched rows on healthy days. State the precondition explicitly, that loss is independent of whether the impression was engaged with. If loss were correlated with engagement, this estimator is biased and the orphan count becomes a lower bound only.
- Characterise the loss as random or structured. Compare the distribution of impressions per session on the affected days against baseline: proportional loss shifts the whole distribution, while a failure that truncates a batch produces a distinctive spike of short sessions. Also check rows per hour to find the start and end of the window.
- Rule out a demand explanation. Sessions per member and engagement events per session should be flat if only logging broke. If sessions also fell, part of the impression drop is real and the two effects must be separated before quoting a loss figure.
- Recommend the operational response: mask these three dates in every impression-denominated rate, re-state them only if the upstream can replay, and note that any experiment reading on those dates is compromised for rate metrics but not for member counts.
Follow-up
- How do you decide whether to backfill the partitions or permanently mask the dates, and what does each choice cost downstream?
- An experiment was reading during those three days. Which of its metrics are still usable and which are not?
- What monitor would have caught this within an hour, and what is its false-positive cost?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Saying no well is a senior skill and it is rarely rehearsed. Think of a time you told someone their analysis was not worth doing, or that the experiment could not answer their question at the sample size available. Explain what you offered instead. Refusal without an alternative reads as obstruction rather than judgement.
Share an experience where you had to influence a product roadmap decis…
Share an experience where you had to influence a product roadmap decision using conflicting data insights.
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Recommend holding a ranker that lifts engagement and hides
A ranking change finished a four-week cluster-randomised test on the home feed. Impressions per engaged session rose 3.1 percent and authored interactions per weekly active member rose 0.4 percent. Negative feedback per 1,000 impressions rose 6 percent, concentrated in hide and not_interested; unfollow was flat. Creator reach concentration rose 1.4 points. The product owner has already briefed the launch upward. In ten minutes give your recommendation, the exchange rate you are applying between engagement and quality, and the specific result that would change your mind. You may request two extra cuts of the data.
Approach
- Put both movements on the same base before arguing about them. Negative feedback is denominated per impression, and impressions per engaged session rose 3.1 percent, so absolute negative actions per engaged session rose about 9.3 percent (1.06 times 1.031), not 6 percent. Say that number out loud; it is usually the first thing nobody has computed.
- Ask whether the negative feedback rise is broad or concentrated: report distinct actors per 1,000 impressions beside the event rate. A rise driven by more members hiding is a distribution problem that affects the median viewer; a rise driven by the same members hiding more is a targeting problem in a segment that may be separable.
- Connect the 1.4 point concentration move to the supply-side guardrail rather than treating it as a curiosity. Pull retained reaching creators by arm and the median viewer's negative feedback on impressions from sub-threshold creators. Concentration is the plausible mechanism that pays for the engagement, and it is paid in creator churn that a four-week window barely registers.
- Read the effect by week with the burn-in excluded, not pooled. A 0.4 percent authored-interaction effect that is 1.1 percent in week 1 and 0.1 percent by week 4 is novelty decay, not a lift. State whether the ranking model was frozen for the test; if it retrained on experiment data, the arms are not independent and a pooled estimate is not interpretable either way.
- Deliver the recommendation as an exchange rate the owner can argue with: this buys roughly N additional hides per additional authored interaction at current volume. Then name the falsifier, for example the week-4 authored-interaction effect holding above 0.3 percent with flat concentration and the negative feedback rise confined to a removable segment.
Follow-up
- The owner launches anyway. What do you instrument on day one, and what is your stop rule?
- Negative feedback rate depends on how reachable the hide control is. Did the treatment change any surface affordance, and how would you know?
- If concentration rose, does the cluster randomisation still hold? Whose feeds leaked into whose?
Disagree with a product manager about push volume
A product manager proposes raising the daily push cap from 3 to 5, citing that members who open a push have 2.4 times the qualified-session rate of members who do not. You have fct_notification_decision with holdout_group in ('send','holdout_type','holdout_global') and 28-day incremental estimates by notification_type. The change is scheduled in a week. In ten minutes, write what you would put in the ticket: what the cited ratio measures, what number should decide this, and what version of the change you would support.
Approach
- Name what the 2.4x measures in one line, without jargon: opening a push and starting a session are both produced by the same intent, so the ratio would look similar if the pushes had been replaced by silent no-ops delivered to the same members. It is a description of who opens things, not an effect.
- Bring the replacement number in the same breath, because a disagreement without an alternative reads as obstruction: incremental qualified sessions per member, send arm minus holdout_global arm, 28 days, split by notification_type. Expect social_reciprocal and recommendation types to differ by a lot, which matters because a cap change does not raise them equally.
- Attack the extrapolation, which is the actual flaw in the proposal. Even a positive increment at cap 3 says nothing about cap 5: frequency response is non-linear and, by construction, notifications 4 and 5 are the lowest-scoring ones the ranker had left. The quantity that decides this is the increment of the marginal notification, which only a cap experiment produces.
- Price the cost the proposal leaves out, in measurable terms rather than as a worry: opt-out rate visible as suppression_reason = 'opt_out' growth, delivery permission loss visible in the delivered over sent ratio, both close to irreversible per member, and both slower than a 28-day session read.
- Offer the version you would support so the disagreement resolves into a plan: a randomised cap ramp on a slice, opt-out and delivery-rate guardrails with a pre-registered stop rule, reads at week 4 and week 8, and agreement in advance on what increment justifies the permanent cap.
Follow-up
- The PM says an eight-week ramp misses the quarter. What is the smallest experiment you would still sign?
- Opt-out is rare. How do you get a usable read on it inside the ramp window?
- The incremental estimate comes back positive but very small. What do you actually recommend?
- 01
Share an experience where you had to influence a product roadmap decision using conflicting data insights.
- 02
A ranking change finished a four-week cluster-randomised test on the home feed. Impressions per engaged session rose 3.1 percent and authored interactions per weekly active member rose 0.4 percent. Negative feedback per 1,000 impressions rose 6 percent, concentrated in hide and not_interested; unfollow was flat. Creator reach concentration rose 1.4 points. The product owner has already briefed the launch upward. In ten minutes give your recommendation, the exchange rate you are applying between engagement and quality, and the specific result that would change your mind. You may request two extra cuts of the data.
- 03
A product manager proposes raising the daily push cap from 3 to 5, citing that members who open a push have 2.4 times the qualified-session rate of members who do not. You have fct_notification_decision with holdout_group in ('send','holdout_type','holdout_global') and 28-day incremental estimates by notification_type. The change is scheduled in a week. In ten minutes, write what you would put in the ticket: what the cited ratio measures, what number should decide this, and what version of the change you would support.
Is this an official Nextdoor interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Nextdoor. Rounds and questions reflect what candidates have reported, not a process Nextdoor has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview loop, and how much preparation time should I plan for?
The interview loop is moderately rigorous, balancing foundational technical coding with deep product sense and experimentation case studies. Most candidates benefit from 3 to 4 weeks of dedicated preparation, focusing heavily on advanced SQL window functions, A/B testing edge cases, and structuring ambiguous product metrics questions.
PracHub interview research ↗What differentiates a good candidate from an exceptional one during the onsite?
Exceptional candidates do not just rush to give a formulaic answer; they proactively clarify ambiguity, discuss potential failure modes, and connect technical metrics back to the broader user experience and business mission. They also demonstrate strong cross-functional empathy, explaining how they partner with engineering and product teams to drive real-world impact.
PracHub interview research ↗What is the company culture like for data scientists at Nextdoor?
The culture emphasizes mission-driven community building, cross-functional collaboration, and an AI-first operational environment. Data scientists operate in semi-embedded structures, giving them high visibility and direct influence over core product decisions while maintaining a healthy balance between analysis and execution.
PracHub interview research ↗What is the typical timeline from the initial recruiter screen to a final offer?
The entire process generally spans 3 to 5 weeks from the initial introductory call to the final decision. This includes a recruiter screen, a technical screening round, and a multi-round onsite loop, followed by prompt debriefs and offer discussions orchestrated by a supportive recruiting team.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22