Whatnot · Data Scientist
Updated · 2026-09-22

Whatnot Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Whatnot, you sit at the intersection of product strategy, engineering, and marketplace growth for the largest live shopping platform in North America and Europe. Your role is vital in decoding the complex dynamics of live auctions, peer-to-peer commerce, and community-driven entertainment across diverse categories like fashion, collectibles, and electronics. You do not just crunch numbers in isolation; you serve as a core strategic partner to product managers, engineers, and category leaders, shaping how millions of buyers and sellers discover, transact, and connect every single day.

SQL is seldom the hardest round and is often the one that eliminates people. The working bar is usually window functions, correct deduplication, and joins that do not silently fan out rows, rather than obscure syntax.

Whatnot candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Diagnose whether a market is supply- or demand-constrainedPrice incentives against contribution margin, not gross bookingsDesign switchback tests that survive supply-side interference

36 min read

Practice 17 Data Scientist prompts
7Company bank questionsSnapshot · Sep 23, 2026 PT
12Candidate experiences ↗Read their reports
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Whatnot, you sit at the intersection of product strategy, engineering, and marketplace growth for the largest live shopping platform in North America and Europe. Your role is vital in decoding the complex dynamics of live auctions, peer-to-peer commerce, and community-driven entertainment across diverse categories like fashion, collectibles, and electronics. You do not just crunch numbers in isolation; you serve as a core strategic partner to product managers, engineers, and category leaders, shaping how millions of buyers and sellers discover, transact, and connect every single day.

Your day-to-day impact involves translating ambiguous, open-ended business challenges into rigorous analytical frameworks and actionable insights. Whether you are driving end-to-end analysis of revenue and category performance in emerging international markets, designing sophisticated experimentation frameworks, or defining core marketplace KPIs, your work directly influences company trajectory. You will build scalable data products, automated reporting dashboards, and forward-looking models that guide major strategic decisions, resource allocation, and feature rollouts across the platform.

Operating in this role requires a rare blend of sharp technical execution, strong commercial judgment, and low-ego cross-functional collaboration. Whatnot moves at an extraordinary pace, demanding a balance between analytical rigor and speed of execution. You will need to embrace ambiguity, prioritize high-impact insights over academic perfection, and communicate complex concepts with absolute clarity to technical and non-technical stakeholders alike.

01

Recruiter Screen

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research
02

Hiring Manager Deep-Dive

reported

Underneath the questions about your past work sits a resourcing question. Given four things worth doing and one of you, which gets done and what happens to the rest? Managers ask because that is the daily texture of the job, and because the answer shows whether you rank work by effort or by what it changes. The weak version sorts by personal interest or by whoever asked most insistently. The strong version ties each candidate piece of work to a decision somebody downstream is waiting on, and then names the one you would drop and who you would tell.

What to demonstrate

  • Whether you rank work by the decision it unblocks or by how interesting the method is
  • How you describe a request you declined, and whether you can say who you said it to
  • Whether your sense of how long something takes survives one follow-up question about the messy part
  • How you decide something is good enough to hand over unfinished

How to prepare

  • Write out your current queue and, next to each item, the decision that stays stalled until it lands. Anything with no waiting decision becomes your example of work you would cut
  • Rehearse turning down a plausible stakeholder request out loud, including the smaller alternative you offered instead
  • Have one case where you shipped a rough answer early and one where you refused to, with the reason that separated them
PracHub interview research
03

Technical Screen

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research
04

Virtual Onsite Panel

reported

A loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.

What to demonstrate

  • Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
  • Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
  • Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience

How to prepare

  • Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
  • For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
  • Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
PracHub interview research

12 candidate reports. Individual accounts describe a particular role and hiring cycle.

Software Engineer

Whatnot Software Engineer interview: DSA pre-screen and two technical rounds

Technical Screen → HR Screen

I started with a structured DSA pre-screen that felt similar to a HireVue or Karat flow. I answered a set of coding prompts and then had a recruiter call. That conversation led into two more technical DSA rounds with engineers from Whatnot. Each session had only a couple of problems, and the interviewers focused on how I was thinking, not just whether I reached the correct output. The process fel…

Read full experience
Full Stack Engineer

Whatnot Full Stack Engineer: recruiter call ended before technical depth

Outcome: rejected

My first step was a quick recruiter call focused almost entirely on my projects. I tried to give a one-sentence summary of each team and company so I could cover my experience efficiently, but the interviewer stopped me earlier than I expected and focused on one project in more detail. The conversation ended shortly afterward, and I was declined. What bothered me was that it never went into much…

Read full experience
Software Engineer

Whatnot Software Engineer interview with HireVue and Karat

Other → Technical Screen → HR Screen

My process started with a HireVue-style interview, followed by a Karat technical interview. The technical portion moved from easy to harder LeetCode questions, and the increase in difficulty was noticeable. After the Karat round, I had a short recruiter screening. If I passed that, I would move on to another technical interview with engineers. That phase was described as including more problem-so…

Read full experience
Software Engineer

Whatnot Software Engineer interview with a one-hour LeetCode-style round

Outcome: rejected

I had a one-hour LeetCode-style round with an engineer. The interviewer was present but very quiet while I worked, so there wasn't much back-and-forth or verbal guidance. The session felt more isolated than I expected, and I ended up not moving forward. The unusual silence is what I remember most. It put more pressure on me to drive the problem-solving process on my own. Location: United States.…

Read full experience
Software Engineer

Whatnot Software Engineer interview: four rounds and a collaborative process

Technical Screen → Other

My interview loop had four rounds. It started with a technical screen focused on core software engineering skills, followed by coding, system design, and product sense rounds. The product sense round felt different from the others because it was more of a discussion than a typical question-and-answer session. The interviewers were kind and supportive throughout, and the overall tone felt collabor…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Reading incentive impact without a cell-level holdout

A bonus in one hour or one zone pulls provider hours and consumer orders from adjacent hours and zones rather than creating them, so a before-and-after read on the treated cell counts displaced volume as incremental and can show a positive result for a spend that produced nothing. Only a randomised holdout at the same granularity as the incentive, or a comparison against untreated cells that share the demand shock, separates increment from displacement. Always state incremental orders per incentive dollar, never total orders in treated cells.

02

Denominator drift in per-active-user metrics

Orders per active consumer falls when acquisition succeeds, because new cohorts transact less than tenured ones, so the metric penalises the thing the company is trying to do. A team that optimises it will quietly prefer weaker acquisition. Decompose into cohort size times cohort frequency, or hold the cohort fixed and read frequency by tenure bucket, before drawing any conclusion about engagement.

03

Never asking what decision the analysis will inform

Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.

04

SQL that silently fans out on a one-to-many join

State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

Sessionise a provider heartbeat stream into supply sessions

hard
sessionisationevent streamscumsum keysinterval arithmetic

You are given pings: provider_id, market_id, event_at_utc, status in ('idle','en_route','engaged','offline'), one row per app heartbeat, nominally every 30 seconds but with gaps. Reconstruct the supply-session table. A session starts at the first non-offline ping and ends at an explicit offline ping, at a market change, or when the gap to the next ping exceeds 10 minutes. Produce session_id, provider_id, market_id, online_at, offline_at, online_seconds, engaged_seconds, en_route_seconds, idle_seconds and end_reason, with the three state components summing to online_seconds exactly in integer seconds.

Approach
  1. Sort by (provider_id, event_at_utc), then build a new_session boolean: first ping of a provider, previous status equals 'offline', market_id changed, or the gap from the previous ping exceeds 600 seconds. A cumsum over that boolean is the session key, and it removes any need for a per-provider Python loop.
  2. Attribute duration to intervals, not to pings: each ping owns the seconds until the next ping inside the same session, and the final ping owns a capped 30 seconds. Because the state seconds are the intervals themselves, they sum to online_seconds by construction rather than by a correction step.
  3. Encode the three terminations distinctly. Gap timeout ends at last_ping + 30s with end_reason 'app_background_timeout'; an explicit offline ping ends at that ping with 'manual_offline'; a session with no terminating event before the data ends is 'session_still_open' with offline_at NaT.
  4. On a market change, close the old session at its last ping in the old market and open the new one at the first ping in the new market; the seconds in between belong to neither session, and the output should say so rather than quietly padding one side.
  5. Aggregate with a single groupby on the session key, pivoting the per-interval state into the three second columns, then assert the sum identity and that consecutive sessions for one provider never overlap.
Follow-up
  • A provider is engaged on a 40-minute order and the app backgrounds mid-order. What does your 10-minute rule do to that session, and what does it do to utilisation?
  • Utilisation divides engaged by online. Which of your three end_reason cases biases it most, and in which direction?

Net take and margin per market-month without ledger fan-out

mediumWorked solution
join fan-outunit economicsattribution window

You have orders (order_id, market_id, completed_at_utc, order_status, gross_booking_cents, provider_payout_cents, consumer_incentive_cents, provider_incentive_cents, tip_cents) and ledger (ledger_id, order_id, entry_type in charge, refund, chargeback, incentive, payout, adjustment, processing_fee; amount_cents signed positive into the platform; posted_at_utc; settlement_status). For completed orders, produce a market-month frame with net take, margin before support cost, and the order count. Attribute refunds, chargebacks and processing fees to the order's completion month, not the posting month, and drop ledger rows whose settlement_status is 'pending' or 'failed'.

Approach
  1. Aggregate the ledger to the order grain before touching orders: filter settlement_status, then pivot entry_type into one summed column each. Merging first and summing after multiplies gross_booking_cents by the number of ledger rows on that order.
  2. Merge the pivoted ledger onto orders with validate='m:1' so an unexpected duplicate raises instead of silently inflating every total.
  3. Net take = gross_booking - provider_payout - consumer_incentive - provider_incentive over completed orders only; tips are excluded on both sides because they pass through to the provider and never enter platform revenue.
  4. Add the signed ledger columns rather than subtracting absolute values: refunds, chargebacks, payouts and processing fees are negative under this convention, so margin = net_take + refund + chargeback + processing_fee. Verify the sign on one known order before trusting the aggregate.
  5. Group by market_id and completed_at_utc month, never posting month; that is the whole point of the attribution rule, and it is why a month's margin is not final until the chargeback window closes.
  6. Carry completed_orders in the output so per-order margin can be recomputed downstream without averaging an average.
Worked solution 30 min
  1. led = ledger[~ledger.settlement_status.isin(['pending','failed'])]; wide = led.pivot_table(index='order_id', columns='entry_type', values='amount_cents', aggfunc='sum', fill_value=0).
  2. comp = orders[orders.order_status == 'completed']; m = comp.merge(wide, on='order_id', how='left', validate='m:1').fillna({col: 0 for col in wide.columns}).
  3. net_take = gross_booking_cents - provider_payout_cents - consumer_incentive_cents - provider_incentive_cents.
  4. margin = net_take + refund + chargeback + processing_fee, using the signed ledger columns directly.
  5. Group by market_id and completed_at_utc.dt.to_period('M'), summing net_take and margin and counting order_id.
EXPECTED RESULTOne row per (market_id, completion_month) with net_take_cents, margin_cents and completed_orders, where margin_cents is less than or equal to net_take_cents in every row given that fee, refund and chargeback entries are non-positive.
Follow-up
  • A chargeback posts two months after completion and changes an already-reported month. How do you publish a metric that is not final, and what lag would you quote?
  • Should the denominator be matched orders, completed orders, or completed-and-settled orders? Argue for one and name what it hides.

Trailing 30-day prior-order counts without rolling or asof

medium
vectorisationsearchsortedwindow semantics

You are given orders: order_id, consumer_id, completed_at_utc (tz-aware UTC), about two million rows, one row per completed order. For every order, compute how many completed orders the same consumer had in the 30 days before that order, counting the window as [t - 30 days, t) so the order itself and any exact-timestamp twin are excluded. You may not use groupby().rolling, merge_asof, or apply over groups. Return the input frame, in its original row order and index, with one added integer column prior_30d.

Approach
  1. Sort once by (consumer_id, completed_at_utc) while keeping the original index, and move to NumPy int64 nanoseconds; the whole problem is two searchsorted calls per group, and a Python loop over two million rows is what makes this fail on time rather than on logic.
  2. Find group boundaries with np.flatnonzero on a consumer_id change mask instead of iterating a groupby object, then slice the timestamp array per block.
  3. Within a block, prior_30d[i] = searchsorted(ts_block, t_i, 'left') - searchsorted(ts_block, t_i - 30 days, 'left'), which is exactly the half-open window and needs no special case for the first order.
  4. State the tie rule out loud: side='left' on the upper bound means simultaneous orders do not count each other, which is the defensible choice when the timestamp has second resolution.
  5. Scatter the result back through the sort permutation so the added column aligns with the caller's frame, and assert the index is unchanged before returning.
Follow-up
  • How does the implementation change if the count must be restricted to the same market?
  • This column will feed a model scored at request time. What leakage would you check for, and which timestamp defines the cut-off?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.

Describe a situation where you had to balance analytical rigor with ex…

medium
behavioural and stakeholder questions

Describe a situation where you had to balance analytical rigor with extreme business urgency; how did you decide when to ship versus when to iterate?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Close with what you would do differently, concretely.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Defend a finding that a provider bonus is incremental but uneconomic

medium
incremental spendholdoutstakeholder conflict

A provider bonus ran in 40 treated market-hours. The program owner's read is +12% completed orders versus the prior week in the treated cells. Your holdout analysis, against untreated cells that shared the same demand shock, puts the effect at 0.06 incremental completed orders per incentive dollar, 95% interval [0.01, 0.11], against contribution margin of $1.40 per completed order. The owner presents the +12% at a review tomorrow and has not seen your number. Deliverable: what you do before the meeting, what you say in it, and the evidence you bring.

Approach
  1. Reproduce the owner's +12% exactly, from their cells and their window, before doing anything else; a disagreement where you cannot reproduce the other number is a credibility fight rather than a measurement one.
  2. Be precise about what you found: the interval [0.01, 0.11] excludes zero, so the bonus does buy orders and you are not claiming otherwise. The argument is about price, not existence, and saying 'it did nothing' hands the owner a refutation built from your own numbers.
  3. Set the bar in the same units as the estimate: breaking even at $1.40 of contribution margin needs 1 / 1.40, about 0.71 incremental orders per dollar. The upper end of your interval, 0.11, is about six and a half times short of that, so every value the data support loses money - which is what makes a wide interval decision-ready without being narrowed.
  4. Invert the ratio, because per-dollar increment is easy to nod at and hard to feel: 0.06 orders per dollar means about $16.70 of bonus per incremental order against $1.40 of margin, the optimistic end 0.11 is about $9.10, and the pessimistic end 0.01 is about $100.
  5. Show the displacement rather than asserting it, and keep the evidence cells out of the comparison set: adjacent untreated hours and adjacent zones where completed orders fell while treated cells rose is the signature of volume moved rather than created. State the precondition - if displacement also reached the difference-in-differences comparison cells, their post-period is depressed by the treatment, the estimate is biased upward, and 0.06 is an upper bound rather than a point estimate.
  6. Separate the measurement question from the decision question and name what would change your mind: a randomised holdout at the same granularity as the incentive, sized in advance, with a date. Offer to run it rather than only to block the program.
Follow-up
  • The owner argues the bonus buys provider retention rather than orders - how would you test that claim, and over what horizon?
  • At what contribution margin per completed order would 0.06 orders per dollar break even, and is that margin reachable in this marketplace?
  • How would you tell displacement across hours apart from a genuine demand shift?

Handle a request to re-cut a test after the readout

hard
pre-registrationselectionguardrails

A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.

Approach
  1. Sort the three requests by one test: could this have been specified before anyone saw the result, and is it symmetric across arms. That test, not the sponsor's seniority, decides what you run.
  2. Decline the gross-bookings switch on mechanism rather than on process: the credit operates by spending incentive dollars, and gross bookings excludes incentive spend by construction, so it cannot see the cost the guardrail exists to catch.
  3. Decline the seven-day window because the credit's payback horizon is longer than the window, so a short read measures the redemption spike rather than the behaviour change, and because the window was chosen after the four-week result was known.
  4. Run the outage exclusion only under stated conditions: the outage is visible in a metric nobody selected, such as requests per market-hour in fct_request, it hit both arms in the same proportion, and it is timestamped independently of this test. Report it as a sensitivity beside the primary, never as a replacement.
  5. Put all three requests in the readout with their status and reasoning, which makes the selection visible and removes the incentive to ask again quietly; then give the sponsor a real path forward: the incentive level at which the credit would break even on contribution margin, and a powered follow-up if that level is reachable.
Follow-up
  • The sponsor says the guardrail was the wrong metric all along - how do you respond?
  • What if the outage is real but hit only the treatment arm?
  • How would you have pre-registered exclusions so that this conversation never happened?
  • 01

    Describe a situation where you had to balance analytical rigor with extreme business urgency; how did you decide when to ship versus when to iterate?

  • 02

    A provider bonus ran in 40 treated market-hours. The program owner's read is +12% completed orders versus the prior week in the treated cells. Your holdout analysis, against untreated cells that shared the same demand shock, puts the effect at 0.06 incremental completed orders per incentive dollar, 95% interval [0.01, 0.11], against contribution margin of $1.40 per completed order. The owner presents the +12% at a review tomorrow and has not seen your number. Deliverable: what you do before the meeting, what you say in it, and the evidence you bring.

  • 03

    A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.

PracHub interview preparation framework
Is this an official Whatnot interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Whatnot. Rounds and questions reflect what candidates have reported, not a process Whatnot has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research
How rigorous is the technical coding portion of the interview?

The technical screen focuses heavily on SQL and occasionally Python, testing your ability to write clean, efficient queries under tight time limits. Expect practical data manipulation challenges involving aggregations, joins, and window functions rather than obscure algorithmic puzzles.

PracHub interview research
What is the typical interview timeline from application to final decision?

The process moves remarkably fast, often wrapping up within two to three weeks from initial recruiter contact to final offer. However, candidates should maintain flexibility, as schedules can occasionally shift based on leadership availability and internal team prioritization.

PracHub interview research
How important is marketplace experience for this role?

While direct experience in e-commerce or two-sided marketplaces is a strong plus, it is not an absolute prerequisite. Interviewers place much higher value on your general product intuition, rigorous statistical thinking, and ability to reason about complex, multi-sided system dynamics.

PracHub interview research
What does the culture of the data team feel like?

The team values low-ego collaboration, high ownership, and extreme speed of execution. Data scientists at Whatnot are expected to be proactive business partners who build practical tools that teams actually use, rather than theoretical models that sit on a shelf.

PracHub interview research
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.