As a Data Scientist at Whatnot, you sit at the intersection of product strategy, engineering, and marketplace growth for the largest live shopping platform in North America and Europe. Your role is vital in decoding the complex dynamics of live auctions, peer-to-peer commerce, and community-driven entertainment across diverse categories like fashion, collectibles, and electronics. You do not just crunch numbers in isolation; you serve as a core strategic partner to product managers, engineers, and category leaders, shaping how millions of buyers and sellers discover, transact, and connect every single day.
Your day-to-day impact involves translating ambiguous, open-ended business challenges into rigorous analytical frameworks and actionable insights. Whether you are driving end-to-end analysis of revenue and category performance in emerging international markets, designing sophisticated experimentation frameworks, or defining core marketplace KPIs, your work directly influences company trajectory. You will build scalable data products, automated reporting dashboards, and forward-looking models that guide major strategic decisions, resource allocation, and feature rollouts across the platform.
Operating in this role requires a rare blend of sharp technical execution, strong commercial judgment, and low-ego cross-functional collaboration. Whatnot moves at an extraordinary pace, demanding a balance between analytical rigor and speed of execution. You will need to embrace ambiguity, prioritize high-impact insights over academic perfection, and communicate complex concepts with absolute clarity to technical and non-technical stakeholders alike.
Recruiter Screen
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Hiring Manager Deep-Dive
reportedUnderneath the questions about your past work sits a resourcing question. Given four things worth doing and one of you, which gets done and what happens to the rest? Managers ask because that is the daily texture of the job, and because the answer shows whether you rank work by effort or by what it changes. The weak version sorts by personal interest or by whoever asked most insistently. The strong version ties each candidate piece of work to a decision somebody downstream is waiting on, and then names the one you would drop and who you would tell.
What to demonstrate
- Whether you rank work by the decision it unblocks or by how interesting the method is
- How you describe a request you declined, and whether you can say who you said it to
- Whether your sense of how long something takes survives one follow-up question about the messy part
- How you decide something is good enough to hand over unfinished
How to prepare
- Write out your current queue and, next to each item, the decision that stays stalled until it lands. Anything with no waiting decision becomes your example of work you would cut
- Rehearse turning down a plausible stakeholder request out loud, including the smaller alternative you offered instead
- Have one case where you shipped a rough answer early and one where you refused to, with the reason that separated them
Technical Screen
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Virtual Onsite Panel
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
12 candidate reports. Individual accounts describe a particular role and hiring cycle.
Whatnot Software Engineer interview: DSA pre-screen and two technical rounds
I started with a structured DSA pre-screen that felt similar to a HireVue or Karat flow. I answered a set of coding prompts and then had a recruiter call. That conversation led into two more technical DSA rounds with engineers from Whatnot. Each session had only a couple of problems, and the interviewers focused on how I was thinking, not just whether I reached the correct output. The process fel…
Read full experienceWhatnot Full Stack Engineer: recruiter call ended before technical depth
My first step was a quick recruiter call focused almost entirely on my projects. I tried to give a one-sentence summary of each team and company so I could cover my experience efficiently, but the interviewer stopped me earlier than I expected and focused on one project in more detail. The conversation ended shortly afterward, and I was declined. What bothered me was that it never went into much…
Read full experienceWhatnot Software Engineer interview with HireVue and Karat
My process started with a HireVue-style interview, followed by a Karat technical interview. The technical portion moved from easy to harder LeetCode questions, and the increase in difficulty was noticeable. After the Karat round, I had a short recruiter screening. If I passed that, I would move on to another technical interview with engineers. That phase was described as including more problem-so…
Read full experienceWhatnot Software Engineer interview with a one-hour LeetCode-style round
I had a one-hour LeetCode-style round with an engineer. The interviewer was present but very quiet while I worked, so there wasn't much back-and-forth or verbal guidance. The session felt more isolated than I expected, and I ended up not moving forward. The unusual silence is what I remember most. It put more pressure on me to drive the problem-solving process on my own. Location: United States.…
Read full experienceWhatnot Software Engineer interview: four rounds and a collaborative process
My interview loop had four rounds. It started with a technical screen focused on core software engineering skills, followed by coding, system design, and product sense rounds. The product sense round felt different from the others because it was more of a discussion than a typical question-and-answer session. The interviewers were kind and supportive throughout, and the overall tone felt collabor…
Read full experiencePracHub editorial advice for the preparation topics above.
Reading incentive impact without a cell-level holdout
A bonus in one hour or one zone pulls provider hours and consumer orders from adjacent hours and zones rather than creating them, so a before-and-after read on the treated cell counts displaced volume as incremental and can show a positive result for a spend that produced nothing. Only a randomised holdout at the same granularity as the incentive, or a comparison against untreated cells that share the demand shock, separates increment from displacement. Always state incremental orders per incentive dollar, never total orders in treated cells.
Denominator drift in per-active-user metrics
Orders per active consumer falls when acquisition succeeds, because new cohorts transact less than tenured ones, so the metric penalises the thing the company is trying to do. A team that optimises it will quietly prefer weaker acquisition. Decompose into cohort size times cohort frequency, or hold the cohort fixed and read frequency by tenure bucket, before drawing any conclusion about engagement.
Never asking what decision the analysis will inform
Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.
SQL that silently fans out on a one-to-many join
State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Sessionise a provider heartbeat stream into supply sessions
You are given pings: provider_id, market_id, event_at_utc, status in ('idle','en_route','engaged','offline'), one row per app heartbeat, nominally every 30 seconds but with gaps. Reconstruct the supply-session table. A session starts at the first non-offline ping and ends at an explicit offline ping, at a market change, or when the gap to the next ping exceeds 10 minutes. Produce session_id, provider_id, market_id, online_at, offline_at, online_seconds, engaged_seconds, en_route_seconds, idle_seconds and end_reason, with the three state components summing to online_seconds exactly in integer seconds.
Approach
- Sort by (provider_id, event_at_utc), then build a new_session boolean: first ping of a provider, previous status equals 'offline', market_id changed, or the gap from the previous ping exceeds 600 seconds. A cumsum over that boolean is the session key, and it removes any need for a per-provider Python loop.
- Attribute duration to intervals, not to pings: each ping owns the seconds until the next ping inside the same session, and the final ping owns a capped 30 seconds. Because the state seconds are the intervals themselves, they sum to online_seconds by construction rather than by a correction step.
- Encode the three terminations distinctly. Gap timeout ends at last_ping + 30s with end_reason 'app_background_timeout'; an explicit offline ping ends at that ping with 'manual_offline'; a session with no terminating event before the data ends is 'session_still_open' with offline_at NaT.
- On a market change, close the old session at its last ping in the old market and open the new one at the first ping in the new market; the seconds in between belong to neither session, and the output should say so rather than quietly padding one side.
- Aggregate with a single groupby on the session key, pivoting the per-interval state into the three second columns, then assert the sum identity and that consecutive sessions for one provider never overlap.
Follow-up
- A provider is engaged on a 40-minute order and the app backgrounds mid-order. What does your 10-minute rule do to that session, and what does it do to utilisation?
- Utilisation divides engaged by online. Which of your three end_reason cases biases it most, and in which direction?
Net take and margin per market-month without ledger fan-out
You have orders (order_id, market_id, completed_at_utc, order_status, gross_booking_cents, provider_payout_cents, consumer_incentive_cents, provider_incentive_cents, tip_cents) and ledger (ledger_id, order_id, entry_type in charge, refund, chargeback, incentive, payout, adjustment, processing_fee; amount_cents signed positive into the platform; posted_at_utc; settlement_status). For completed orders, produce a market-month frame with net take, margin before support cost, and the order count. Attribute refunds, chargebacks and processing fees to the order's completion month, not the posting month, and drop ledger rows whose settlement_status is 'pending' or 'failed'.
Approach
- Aggregate the ledger to the order grain before touching orders: filter settlement_status, then pivot entry_type into one summed column each. Merging first and summing after multiplies gross_booking_cents by the number of ledger rows on that order.
- Merge the pivoted ledger onto orders with validate='m:1' so an unexpected duplicate raises instead of silently inflating every total.
- Net take = gross_booking - provider_payout - consumer_incentive - provider_incentive over completed orders only; tips are excluded on both sides because they pass through to the provider and never enter platform revenue.
- Add the signed ledger columns rather than subtracting absolute values: refunds, chargebacks, payouts and processing fees are negative under this convention, so margin = net_take + refund + chargeback + processing_fee. Verify the sign on one known order before trusting the aggregate.
- Group by market_id and completed_at_utc month, never posting month; that is the whole point of the attribution rule, and it is why a month's margin is not final until the chargeback window closes.
- Carry completed_orders in the output so per-order margin can be recomputed downstream without averaging an average.
Worked solution 30 min
- led = ledger[~ledger.settlement_status.isin(['pending','failed'])]; wide = led.pivot_table(index='order_id', columns='entry_type', values='amount_cents', aggfunc='sum', fill_value=0).
- comp = orders[orders.order_status == 'completed']; m = comp.merge(wide, on='order_id', how='left', validate='m:1').fillna({col: 0 for col in wide.columns}).
- net_take = gross_booking_cents - provider_payout_cents - consumer_incentive_cents - provider_incentive_cents.
- margin = net_take + refund + chargeback + processing_fee, using the signed ledger columns directly.
- Group by market_id and completed_at_utc.dt.to_period('M'), summing net_take and margin and counting order_id.
Follow-up
- A chargeback posts two months after completion and changes an already-reported month. How do you publish a metric that is not final, and what lag would you quote?
- Should the denominator be matched orders, completed orders, or completed-and-settled orders? Argue for one and name what it hides.
Trailing 30-day prior-order counts without rolling or asof
You are given orders: order_id, consumer_id, completed_at_utc (tz-aware UTC), about two million rows, one row per completed order. For every order, compute how many completed orders the same consumer had in the 30 days before that order, counting the window as [t - 30 days, t) so the order itself and any exact-timestamp twin are excluded. You may not use groupby().rolling, merge_asof, or apply over groups. Return the input frame, in its original row order and index, with one added integer column prior_30d.
Approach
- Sort once by (consumer_id, completed_at_utc) while keeping the original index, and move to NumPy int64 nanoseconds; the whole problem is two searchsorted calls per group, and a Python loop over two million rows is what makes this fail on time rather than on logic.
- Find group boundaries with np.flatnonzero on a consumer_id change mask instead of iterating a groupby object, then slice the timestamp array per block.
- Within a block, prior_30d[i] = searchsorted(ts_block, t_i, 'left') - searchsorted(ts_block, t_i - 30 days, 'left'), which is exactly the half-open window and needs no special case for the first order.
- State the tie rule out loud: side='left' on the upper bound means simultaneous orders do not count each other, which is the defensible choice when the timestamp has second resolution.
- Scatter the result back through the sort permutation so the added column aligns with the caller's frame, and assert the index is unchanged before returning.
Follow-up
- How does the implementation change if the count must be restricted to the same market?
- This column will feed a model scored at request time. What leakage would you check for, and which timestamp defines the cut-off?
Optimize a slow-running query that joins multi-terabyte event logs for…
Optimize a slow-running query that joins multi-terabyte event logs for user streaming sessions and checkout funnels.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Extract and aggregate daily buyer conversion rates across multiple geo…
Extract and aggregate daily buyer conversion rates across multiple geographic regions, handling null values and edge cases gracefully.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Approved providers who never matched and never went online
dim_user has user_id, side, provider_status, provider_approved_at_utc. fct_request has request_id, provider_id and matched_at_utc, where provider_id is NULL on every request that never matched. fct_supply_session has session_id, provider_id, online_at_utc. List providers whose provider_status is 'approved' and whose provider_approved_at_utc falls in a given calendar month, who have never appeared as fct_request.provider_id and who have no row in fct_supply_session. Return user_id, provider_approved_at_utc and days since approval, oldest first.
Approach
- Use NOT EXISTS for both exclusions, correlated on provider_id. NOT EXISTS is unaffected by NULLs in the subquery and lets the planner stop at the first matching row.
- If you reach for NOT IN, provider_id is nullable, so the subquery result contains at least one NULL, every comparison evaluates to UNKNOWN instead of TRUE, and the query returns zero rows. There is no error and no warning; an empty result set reads as good news.
- The correct alternatives are a LEFT JOIN with WHERE right_key IS NULL, or NOT IN with an explicit WHERE provider_id IS NOT NULL inside the subquery. Both work; NOT EXISTS is the one that survives someone later making a second column nullable.
- Filter the cohort on provider_approved_at_utc within the month and provider_status = 'approved', so suspended and deactivated accounts do not inflate what will be read as an onboarding failure.
- Compute days since approval as a date difference on the approval timestamp, and sanity check the result size against the cohort size before drawing any conclusion.
Worked solution 20 min
- Run SELECT COUNT(*) FROM fct_request WHERE provider_id IS NULL. A non-zero count is the proof that NOT IN is unusable against this column.
- Write the cohort CTE and count it, so the denominator is known before any filtering.
- Add the two NOT EXISTS clauses one at a time and record the row count after each; which one does most of the work is itself the finding.
- Order by provider_approved_at_utc ascending and add the days-since column.
Follow-up
- The list is 40 percent of the approval cohort. Is that an onboarding failure or a data problem, and which single query settles it?
- How would you separate providers who never went online from providers who went online and were never offered anything, given that fct_request only records the provider who actually matched?
How do you manage competing priorities when multiple product teams req…
How do you manage competing priorities when multiple product teams request custom analytics support simultaneously?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
Whatnot just launched a feature allowing buyers to tip sellers during …
Whatnot just launched a feature allowing buyers to tip sellers during live auctions; how would you evaluate its success?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How do you balance conflicting marketplace metrics, such as buyer acqu…
How do you balance conflicting marketplace metrics, such as buyer acquisition growth versus seller retention and churn?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design a core metric framework to measure the health and…
How would you design a core metric framework to measure the health and engagement of a new live-shopping category on the platform?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How do you determine statistical significance when dealing with highly…
How do you determine statistical significance when dealing with highly skewed transaction value distributions and low-frequency buyers?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
Explain how you would approach causal inference and impact measurement…
Explain how you would approach causal inference and impact measurement when a randomized controlled trial is practically impossible to execute.
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
- State the primary metric and the minimum effect worth shipping, then size the test.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Estimate a single-market fee change without a control group
Regulation forces one market to cut its service fee on a known date. Randomisation is impossible and only that market is affected. You have 52 weeks of pre-period weekly series for every market: completed orders, active consumers, SLA fill rate, provider online hours and net take rate. Estimate the effect on completed orders per active consumer over the 12 weeks after the change. Specify the estimator, the donor pool and its exclusions, the pre-period fit criterion you would accept, and how you would produce a p-value with exactly one treated unit.
Approach
- Choose synthetic control as the estimator. Fit non-negative donor weights that sum to one, matching the treated market pre-period outcome path plus a small set of predictors such as vertical, launch age, price level and supply density. The convexity constraint is the point: it forbids extrapolation beyond the donor support, which an unrestricted regression fit would happily do.
- Build the donor pool by exclusion, not by convenience. Drop markets with their own policy or pricing change inside the window, and drop markets adjacent enough that providers or consumers can cross into the treated market within a session. A spilled-over donor is contaminated toward the treated path and shrinks the estimated effect toward zero.
- Set the fit criterion before looking at the post period. Require a small pre-period RMSPE relative to the effect size that would change the decision, and validate it out of sample by fitting weights on weeks 1 to 40 and checking fit on weeks 41 to 52, so the weights are not tuned to pre-period noise.
- Produce inference by permutation, because there is one treated unit and no cluster to compute a standard error over. Run the identical procedure treating each donor as if it were treated, compute the ratio of post-period to pre-period RMSPE for each, and rank the true treated market. With 30 donors the smallest attainable p-value is 1/31 = 0.032, so state that floor rather than reporting a conventional significance claim.
- Add a placebo in time: move the treatment date 12 weeks earlier and confirm the method produces a gap near zero. A method that manufactures an effect on a date when nothing happened cannot be trusted on the date when something did.
- Report the ratio decomposed into its parts. A fee cut moves completed orders and active consumers together, so completed orders per active consumer can stay flat while both components move substantially, and the decomposition is what the decision actually needs.
Worked solution 35 min
- Assemble the weekly panel and screen the donor pool for policy changes and geographic spillover, recording each exclusion and its reason.
- Fit donor weights on pre-period weeks 1 to 40 and validate the fit on weeks 41 to 52, reporting pre-period RMSPE.
- Compute the post-period gap between the treated market and its synthetic counterpart for each of the 12 weeks.
- Run placebo-in-space over every donor, rank the treated post/pre RMSPE ratio, and convert the rank into a permutation p-value with its 1/(N+1) floor.
- Run placebo-in-time at a fake treatment date 12 weeks early and confirm a near-zero gap.
- Decompose the ratio into completed orders and active consumers and report both alongside the ratio.
Follow-up
- Three markets get the same regulation on three different dates. What changes in the estimator, and what goes wrong with a naive two-way fixed-effects specification?
- Pre-period fit is excellent but the treated market is the largest in the pool. What should you suspect about the weights?
- How would you separate the effect of the fee change from the effect of the consumer-facing price change it caused?
Requests fell four percent on one local day only
Daily requests in one market fell 4% against the previous day and recovered the day after. The dashboard groups fct_request by DATE(requested_at_local), and dim_market.timezone holds the market's IANA zone. Before anyone opens an incident, rule out calendar and clock causes. Deliverable: the ordered checks you run, the arithmetic for the largest mechanical cause you find, and the residual threshold at which you would still escalate.
Approach
- Compare like weekday to like weekday. Marketplace demand has a strong weekly cycle, so a day-over-day comparison that crosses a weekend boundary is not a signal at all. Use the same weekday for the trailing four to eight weeks as the baseline.
- Read the day's length from the zone, not from the row labels. requested_at_local is a naive wall-clock timestamp, so on an autumn fall-back day the repeated hour lands on the same label twice: COUNT(DISTINCT date_trunc('hour', requested_at_local)) returns 24 for a 25-hour day, never 25. That count can only ever catch spring-forward, and even then only in a market busy enough that every surviving hour has at least one request. Compute the length from dim_market.timezone instead - ((d + 1)::timestamp AT TIME ZONE m.timezone) - (d::timestamp AT TIME ZONE m.timezone) returns 23, 24 or 25 hours for local date d, does not depend on volume, and still returns 23 or 25 in the zones whose transition falls at local midnight, where the non-existent instant resolves forward.
- Size the clock effect properly. The shortfall is not one hour in 24, which would be 4.2%, but the share of daily requests that normally falls in the hour that disappeared, read off the trailing same-weekday hourly profile. In most zones the transition lands overnight, so the true mechanical effect is usually well under one percent. A fall-back day runs the other way, and its data-side fingerprint is the duplicated label carrying roughly twice its usual share - which is the only way that case shows up at all if you are counting distinct local hours.
- Check the local calendar for the date: public holidays, school terms, large scheduled events, severe weather. These move the whole day rather than a single hour, which is how you tell them apart from a clock change.
- Escalate only the residual. Subtract the weekday effect and the clock effect, then express what is left in units of the same-weekday standard deviation over the trailing eight occurrences, and give the threshold you are using.
Follow-up
- The same 4% drop lands on the same local date in five markets across three timezones. What do you check first now?
- The product dashboard groups by local date and the finance report by UTC date. What reconciles the two?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.
Describe a situation where you had to balance analytical rigor with ex…
Describe a situation where you had to balance analytical rigor with extreme business urgency; how did you decide when to ship versus when to iterate?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Close with what you would do differently, concretely.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- What did you decide not to do, and why?
- How did you know the outcome was caused by your change?
Defend a finding that a provider bonus is incremental but uneconomic
A provider bonus ran in 40 treated market-hours. The program owner's read is +12% completed orders versus the prior week in the treated cells. Your holdout analysis, against untreated cells that shared the same demand shock, puts the effect at 0.06 incremental completed orders per incentive dollar, 95% interval [0.01, 0.11], against contribution margin of $1.40 per completed order. The owner presents the +12% at a review tomorrow and has not seen your number. Deliverable: what you do before the meeting, what you say in it, and the evidence you bring.
Approach
- Reproduce the owner's +12% exactly, from their cells and their window, before doing anything else; a disagreement where you cannot reproduce the other number is a credibility fight rather than a measurement one.
- Be precise about what you found: the interval [0.01, 0.11] excludes zero, so the bonus does buy orders and you are not claiming otherwise. The argument is about price, not existence, and saying 'it did nothing' hands the owner a refutation built from your own numbers.
- Set the bar in the same units as the estimate: breaking even at $1.40 of contribution margin needs 1 / 1.40, about 0.71 incremental orders per dollar. The upper end of your interval, 0.11, is about six and a half times short of that, so every value the data support loses money - which is what makes a wide interval decision-ready without being narrowed.
- Invert the ratio, because per-dollar increment is easy to nod at and hard to feel: 0.06 orders per dollar means about $16.70 of bonus per incremental order against $1.40 of margin, the optimistic end 0.11 is about $9.10, and the pessimistic end 0.01 is about $100.
- Show the displacement rather than asserting it, and keep the evidence cells out of the comparison set: adjacent untreated hours and adjacent zones where completed orders fell while treated cells rose is the signature of volume moved rather than created. State the precondition - if displacement also reached the difference-in-differences comparison cells, their post-period is depressed by the treatment, the estimate is biased upward, and 0.06 is an upper bound rather than a point estimate.
- Separate the measurement question from the decision question and name what would change your mind: a randomised holdout at the same granularity as the incentive, sized in advance, with a date. Offer to run it rather than only to block the program.
Follow-up
- The owner argues the bonus buys provider retention rather than orders - how would you test that claim, and over what horizon?
- At what contribution margin per completed order would 0.06 orders per dollar break even, and is that margin reachable in this marketplace?
- How would you tell displacement across hours apart from a genuine demand shift?
Handle a request to re-cut a test after the readout
A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.
Approach
- Sort the three requests by one test: could this have been specified before anyone saw the result, and is it symmetric across arms. That test, not the sponsor's seniority, decides what you run.
- Decline the gross-bookings switch on mechanism rather than on process: the credit operates by spending incentive dollars, and gross bookings excludes incentive spend by construction, so it cannot see the cost the guardrail exists to catch.
- Decline the seven-day window because the credit's payback horizon is longer than the window, so a short read measures the redemption spike rather than the behaviour change, and because the window was chosen after the four-week result was known.
- Run the outage exclusion only under stated conditions: the outage is visible in a metric nobody selected, such as requests per market-hour in fct_request, it hit both arms in the same proportion, and it is timestamped independently of this test. Report it as a sensitivity beside the primary, never as a replacement.
- Put all three requests in the readout with their status and reasoning, which makes the selection visible and removes the incentive to ask again quietly; then give the sponsor a real path forward: the incentive level at which the credit would break even on contribution margin, and a powered follow-up if that level is reachable.
Follow-up
- The sponsor says the guardrail was the wrong metric all along - how do you respond?
- What if the outage is real but hit only the treatment arm?
- How would you have pre-registered exclusions so that this conversation never happened?
- 01
Describe a situation where you had to balance analytical rigor with extreme business urgency; how did you decide when to ship versus when to iterate?
- 02
A provider bonus ran in 40 treated market-hours. The program owner's read is +12% completed orders versus the prior week in the treated cells. Your holdout analysis, against untreated cells that shared the same demand shock, puts the effect at 0.06 incremental completed orders per incentive dollar, 95% interval [0.01, 0.11], against contribution margin of $1.40 per completed order. The owner presents the +12% at a review tomorrow and has not seen your number. Deliverable: what you do before the meeting, what you say in it, and the evidence you bring.
- 03
A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.
Is this an official Whatnot interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Whatnot. Rounds and questions reflect what candidates have reported, not a process Whatnot has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How rigorous is the technical coding portion of the interview?
The technical screen focuses heavily on SQL and occasionally Python, testing your ability to write clean, efficient queries under tight time limits. Expect practical data manipulation challenges involving aggregations, joins, and window functions rather than obscure algorithmic puzzles.
PracHub interview research ↗What is the typical interview timeline from application to final decision?
The process moves remarkably fast, often wrapping up within two to three weeks from initial recruiter contact to final offer. However, candidates should maintain flexibility, as schedules can occasionally shift based on leadership availability and internal team prioritization.
PracHub interview research ↗How important is marketplace experience for this role?
While direct experience in e-commerce or two-sided marketplaces is a strong plus, it is not an absolute prerequisite. Interviewers place much higher value on your general product intuition, rigorous statistical thinking, and ability to reason about complex, multi-sided system dynamics.
PracHub interview research ↗What does the culture of the data team feel like?
The team values low-ego collaboration, high ownership, and extreme speed of execution. Data scientists at Whatnot are expected to be proactive business partners who build practical tools that teams actually use, rather than theoretical models that sit on a shelf.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22