Waymo · Data Scientist
Updated · 2026-09-24

Waymo Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Waymo, you sit at the intersection of complex autonomous vehicle engineering, advanced machine learning, and commercial ride-hailing operations. Your primary mission is to help the company make the most informed, data-driven decisions while scaling the Waymo Driver—the world's most experienced driver—across new geographies, vehicle platforms, and weather conditions. Autonomous driving presents a fundamentally unique paradigm in data science, moving far beyond traditional web analytics into dense simulation data, rare event rate estimation, and rigorous on-road performance evaluation.

Allocate prep to your weakest link rather than your favourite topic. Of the three things that usually gate the outcome (SQL that is correct under messy joins, sound reasoning about experiments, and structured framing of an open-ended problem), candidates tend to over-invest in modelling theory and under-invest in framing.

Waymo candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Size safety stock under variable lead timesSeparate true demand from censored, stocked-out salesPrice a service-level change in working capital

37 min read

Practice 17 Data Scientist prompts
33Company bank questionsSnapshot · Sep 28, 2026 PT
12Candidate experiences ↗Read their reports
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Waymo, you sit at the intersection of complex autonomous vehicle engineering, advanced machine learning, and commercial ride-hailing operations. Your primary mission is to help the company make the most informed, data-driven decisions while scaling the Waymo Driver—the world's most experienced driver—across new geographies, vehicle platforms, and weather conditions. Autonomous driving presents a fundamentally unique paradigm in data science, moving far beyond traditional web analytics into dense simulation data, rare event rate estimation, and rigorous on-road performance evaluation.

Your work directly influences whether software updates, new safety protocols, and large-scale infrastructure footprints are ready for public deployment. Whether you are building incrementality measurement frameworks for digital media, modeling weather patterns and their impact on fleet safety, or optimizing multi-city fleet orchestration, you collaborate hand-in-hand with engineering, product, and operations teams. You will tackle deeply ambiguous problems by scoping technical priorities, establishing novel statistical methodologies, and translating complex data signals into actionable strategies for senior leadership.

Expect an environment that is deeply data-driven, intellectually rigorous, and fast-paced. You will be expected to balance scientific depth with business pragmatism, ensuring that every metric and evaluation framework you design directly supports Waymo's core mission of safety, compliance, and commercial scale. Succeeding in this role requires a rare blend of advanced statistical intuition, robust programming capabilities, and the cross-functional communication skills needed to drive alignment across diverse technical teams.

01

Recruiter Screen

reported

Most candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.

What to demonstrate

  • Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
  • Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
  • The substance of the questions you ask back, which an experienced screener reads as a level signal

How to prepare

  • Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
  • Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
  • Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
PracHub interview research ↗
02

Technical Screen

reported

A handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.

What to demonstrate

  • Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
  • Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
  • Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.

How to prepare

  • Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
  • Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
  • If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
PracHub interview research ↗
03

Onsite Loop

reported

Where a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.

What to demonstrate

  • Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
  • Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
  • Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
  • Whether you ask what decision is actually on the table before explaining anything

How to prepare

  • Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
  • For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
  • Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub interview research ↗
04

Technical Assessments

reported

This round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.

What to demonstrate

  • Whether your row counts survive each join, and whether you notice on your own when they do not
  • Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
  • Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
  • Reaching a defensible answer inside the window instead of a refined one after it

How to prepare

  • Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
  • Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
  • Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
PracHub interview research ↗
05

Behavioral Evaluations

reported

Most of the weight in this round sits on the disagreement questions. Data work routinely produces an answer someone senior did not want, and the interviewer is trying to learn what you do in that hour. Both failure modes are common: folding as soon as a director pushes back, and treating the pushback as ignorance to be corrected with a better chart. A strong answer usually contains a specific thing the other person knew that you did not, and describes how you found out whether it changed the conclusion.

What to demonstrate

  • Whether you can state the other side's argument accurately before you explain why you disagreed
  • What you treated as evidence during the disagreement, such as a rerun under their assumption or a holdout check, rather than persuasion technique
  • Whether you distinguish being overruled from being wrong, and can give an example of each

How to prepare

  • Write out one disagreement where you turned out to be wrong, and say what in the data misled you. Candidates prepare the story where they were right, and the follow-up asks for the other one.
  • For your main disagreement story, be ready to say what result would have made you drop your position. If no such result exists, you were not arguing from the data.
  • Practise stating the opposing position out loud in one sentence the stakeholder would accept, then continue the story.
PracHub interview research ↗

12 candidate reports. Individual accounts describe a particular role and hiring cycle.

Data Scientist

Waymo Data Scientist Interview Experience — Rare Events, Sampling, and Probability Coding

HR Screen → Technical ScreenOutcome: rejected

A recruiter contacted me. After screening and a technical screen, I went into four interview rounds. I've made the original questions less specific for others preparing. Data Intuition, first round A new version performs significantly better in offline evaluation, in simulation, but shows no difference after going live. What could cause that? How would you model a rare event using a log form, and…

Read full experience
Software Engineer

Waymo Software Engineer Interview Experience — Navigation Constraints and Mapping Systems

Technical Screen → Onsite

The author lists a Waymo phone screen followed by onsite system-design and coding interviews. The screen involved a robot moving in four directions, returning to its starting point, and avoiding obstacles. The design discussion asked about a vehicle fleet collecting map information. One onsite coding task matched dictionary words against input containing repeated letters. Another used a map with…

Read full experience
Software Engineer

Waymo Software Engineer Interview Experience — A Simulation Design Round After a Recruiter Reassurance

OnsiteOutcome: rejected

Coding: 317. The second problem was traversing a graph with DFS, fairly simple. There was a follow-up: come up with an algorithm to prove your answer is correct. A strange follow-up. Behavioral: Very standard. A project I'm proud of, how to handle priorities, things like that. System design: This was the first round of the interview. What a disastrous start. I was interviewing with the simulation…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Averaging rates across SKU-locations instead of re-summing

Fill rate, turns, OEE and on-time rate are all ratios whose denominators differ by orders of magnitude between cells, so an unweighted mean gives a slow-moving C item at a small node the same vote as a high-volume A item at a national node. The blended figure then moves whenever the portfolio mix moves, and it can improve in every cell while the company-level ratio worsens, or the reverse, which is Simpson's paradox with a warehouse attached. Always sum numerator and denominator to the reporting level and divide once, and when a rate must be compared across nodes, standardise on a fixed SKU mix before reading anything into the difference.

02

Sizing safety stock as z times sigma_D times the square root of lead time

That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.

03

Over-explaining the method and under-explaining the implication

Lead with the answer and what you would do about it, then give the approach when asked. Roughly one sentence of method per three of implication is the right ratio for a stakeholder-facing answer; the interviewer already knows what a regression is.

04

Analysing at a different unit than the one randomised

Say out loud what was randomised (user, device, account, cluster) and make the analysis unit match, or account for the clustering with cluster-robust standard errors, the delta method, or aggregation up to the randomised unit. Randomising users and then running a test over sessions understates variance and inflates the false-positive rate.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

How do you approach rate estimation with rare events—such as critical …

medium
statistics and probability

How do you approach rate estimation with rare events—such as critical safety interventions—where occurrences are extremely sparse in both real-world and simulation data?

Approach
  1. Sanity-check the answer against a simple bound or a simulated case.
  2. Say what the estimate is of, and over what population it generalises.
  3. Write down the assumption the method needs before you use the method.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • How would you explain this result to someone who does not know statistics?

How would you monitor production machine learning models for data drif…

medium
machine learning and modelling

How would you monitor production machine learning models for data drift and automatically trigger alerts for regression events?

Approach
  1. Check what information would not exist at prediction time, and exclude it.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

Simulate service levels under a variable inbound lead time

mediumWorked solution
safety stockmonte carloservice level

Daily demand for one SKU at one node is Poisson with mean 40. Replenishment lead time is 7 days with probability 0.7, 10 days with 0.2, and 14 days with 0.1, drawn independently per order. The policy is continuous review: order Q = 400 whenever the inventory position drops to the reorder point. Size safety stock two ways for a nominal 95 percent target, once as z times sigma_D times sqrt(mean lead time) and once with the variable-lead-time term. Simulate both and report achieved cycle service level and achieved unit fill rate.

Approach
  1. Compute both safety stock numbers analytically before simulating, so the simulation is a check on arithmetic you already understand rather than the source of the answer.
  2. Track inventory position (on hand plus on order) for the reorder trigger and on hand separately for the stockout test. Triggering off on hand alone re-orders while stock is already in transit and silently changes the policy you are measuring.
  3. Define the two outcomes precisely and separately: cycle service level is the fraction of replenishment cycles containing at least one unit of unmet demand, while fill rate is total units shipped from stock over total units demanded. They answer different questions and z targets only the first.
  4. Before running anything, bound the achievable service level by hand from the lead-time mixture. Each lead-time branch is a separate Poisson demand, so cycle service level is the probability-weighted sum of P(Poisson(40 x L) <= reorder point) across the three branches, and a reorder point that cannot cover the 14-day branch caps the policy at 0.9 no matter how the simulation is tuned.
  5. Run enough cycles that Monte Carlo error is small against the gap you are trying to see. The standard error of a service level is sqrt(p*(1-p)/n_cycles), about 0.005 at p = 0.7 with 10,000 cycles, so the 20-point gap between the two policies is unambiguous and a 1-point one is not.
  6. Let unmet demand be lost rather than backordered, state that choice, and note that switching to full backordering raises measured fill rate for the same policy.
Worked solution 35 min
  1. Compute mean lead time 8.3 days, variance of lead time 5.01, sigma_D = sqrt(40) = 6.32, and mean demand 40 per day.
  2. Naive safety stock: 1.645 * 6.32 * sqrt(8.3) = 30 units, giving a reorder point of 332 + 30 = 362. Correct safety stock: 1.645 * sqrt(8.3 * 40 + 40^2 * 5.01) = 1.645 * 91.4 = 150 units, giving a reorder point of 482.
  3. Write one simulation loop parameterised by safety stock: draw daily Poisson demand, decrement on hand, record shortfall, trigger an order of 400 when position hits the reorder point, and schedule the receipt after a sampled lead time.
  4. Run 10,000 cycles per policy with a burn-in of at least two lead times discarded so the starting position does not flatter the first cycles.
  5. Report achieved cycle service level, achieved fill rate and the average on-hand position for each of the two safety stock settings.
EXPECTED RESULTNaive safety stock is 30 units and the variable-lead-time figure is 150 units, a factor of five. At Q = 400 with lost sales, the naive policy achieves a cycle service level of about 0.70 and the correct policy about 0.90, so the naive policy misses the nominal 0.95 by roughly 25 points and the correct one by 5. The naive ceiling is arithmetic, not simulation noise: with a reorder point of 362 the 7-day branch is essentially always covered, the 10-day branch survives with probability P(Poisson(400) <= 362) = 0.029, and the 14-day branch never does, so cycle service level cannot exceed 0.7 + 0.2*0.029 + 0.1*0 = 0.706, and reorder-point undershoot of about 20 units under daily demand arrivals pulls it to roughly 0.70. The correct policy tops out near 0.90 for the same structural reason: a reorder point of 482 covers the 10-day branch almost completely (P = 0.99997) but not the 14-day one (P = 0.0004), so that 0.1 of probability mass is near-pure loss. Fill rate is higher than cycle service level under both, about 0.92 against 0.70 and about 0.98 against 0.90, because expected units short per cycle (roughly 33 and 10) are small against Q = 400.
Follow-up
  • With Q = 100 against lead-time demand of 332 units, three or four orders are open at once. What happens to the meaning of a replenishment cycle, and to each of the two service measures?
  • Lead time and demand are correlated because the supplier is capacity constrained in peak weeks. What does that do to the formula you used?
  • How would you convert the extra 120 units of safety stock into an annual cost the business can weigh against the service gain?

Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Diagnostic, scored before you study anything
  • Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
  • Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
  • Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.

Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Largest gap: find the boundary rather than the subject
  • Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
  • Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
  • Re-solve one of them from memory four hours later, on paper, with nothing open.

Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Largest gap: drill the blocking sub-skill
  • Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
  • Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
  • Have someone else read your one-sentence rule and find the precondition you left out.

Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Second gap, plus maintenance on your strongest area
  • Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
  • Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
  • Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.

Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05The gap that is not a skill
  • Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
  • Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
  • Deliver one answer again with a hard 90-second limit to force structure before detail.

Deliverable: Two recordings with a counted improvement in time-to-first-question.

Practice prompt ↗Practice prompt ↗
06Retest under day-one conditions
  • Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
  • Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
  • Write which single block you would still lose the offer on.

Deliverable: A second scored rubric placed next to the first, with one named remaining risk.

Practice prompt ↗Practice prompt ↗
07Full loop under interview conditions
  • Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
  • Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
  • Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.

Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.

How do you handle interference and spillover effects between treatment…

medium
behavioural and stakeholder questions

How do you handle interference and spillover effects between treatment and control groups in a localized urban transportation network?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. State the situation in two sentences and spend the rest on your reasoning.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What would you do differently if you ran that project again?
  • What did you decide not to do, and why?

Tell a sponsor the effect cannot be measured in six weeks

hard
powerselection biasquasi experimentssaying no

A VP wants a read in six weeks on a forward-stocking policy now live at eight of forty nodes. Replenishment lead time on the affected lanes is five weeks, the eight nodes were chosen by the ops team because they were the worst performers, and the cross-node standard deviation of weekly fill rate is about 3 points. The VP has already told his leadership team that a number is coming. Tell him what you can and cannot deliver in six weeks, and propose the design and timeline you would commit to instead.

Approach
  1. Do the power arithmetic in front of him rather than asserting the test is underpowered: comparing eight treated against thirty-two untreated node means with a between-node standard deviation of 3 points gives a standard error near 3 times the square root of one eighth plus one thirty-second, about 1.2 points, so the minimum detectable effect at 80 percent power and a two-sided 5 percent test is roughly 3.3 points. Any real effect smaller than that comes back as a null you cannot interpret.
  2. Name the burn-in problem separately: with a five-week lead time, six weeks covers barely one replenishment cycle, so whatever you measure is the transition rather than the new steady state, and the transition usually looks worse than the policy is.
  3. Name the selection problem third: nodes picked because they were worst will improve toward the network mean without any policy, so a simple before-and-after at those nodes is biased upward and will over-claim.
  4. Offer what is genuinely deliverable in six weeks: an implementation read, meaning whether the policy is actually in force at the eight nodes, whether inventory positions moved as designed, and whether any leading indicator such as short_reason_code mix is moving, framed explicitly as operational verification and not an effect estimate.
  5. Propose the real design with dates: extend to sixteen weeks covering roughly three cycles, use a difference in differences against matched comparison nodes chosen on pre-period fill rate and volume, cluster variance at the node, and pre-register the burn-in window that will be excluded.
  6. Help him with the commitment he already made: give him the exact wording for what he reports at week six, so the honest answer arrives as something he can say rather than as a refusal.
Follow-up
  • He asks you to add the remaining thirty-two nodes to the rollout next month. What does that do to your design?
  • If the effect really is 1 point, is the policy worth keeping, and how would you ever know?

Prioritise transport, plant and commercial in one week

medium
prioritisationstakeholder managementdecision value

Three requests land in the same week. Transport wants landed cost per delivered unit split by lane and mode from fct_shipment_leg for a carrier bid that closes in nine days. A plant manager wants OEE on two work centres decomposed into availability, performance and quality from fct_production_run for a capital request due next month. Commercial wants a fill-rate root cause from fct_order_line for an account review on Thursday. You are the only data scientist and have about four working days. Give your order, the rule that produced it, and what you say to the two who wait.

Approach
  1. Rank by the decision each request feeds and by whether your input can still change it, not by seniority or by who asked loudest: a bid closing in nine days is a live decision with a large irreversible spend attached, a capital request due next month has slack, and an account review has a fixed date but a smaller reversible decision.
  2. Check reversibility and blast radius: carrier rates lock for a contract period across every lane, so an error or an absence there is expensive for a year, while the fill-rate story can be revised next week.
  3. Look for the cheap partial that unblocks someone else: the fill-rate cut by node, by short_reason_code and by week is a few hours of work and covers most of what commercial needs on Thursday, so it does not have to wait behind the bid.
  4. Sequence with explicit time boxes: the commercial cut first because it is short and date-locked, the lane and mode cost analysis next with the allocation basis for multi-stop loads stated up front, and the OEE decomposition scheduled into the following week with a date the plant manager can hold you to.
  5. Tell the plant manager directly and early, with a date rather than a maybe, and say what you need from him in the meantime so the delay produces something.
  6. Escalate the trade-off rather than absorbing it silently: your manager should know a capital request slipped a week, because that is a business choice and not yours alone to make.
Follow-up
  • The plant manager escalates and his director asks you to reorder. What do you do?
  • Which of the three would you push back on entirely, and what would you offer instead?
  • 01

    How do you handle interference and spillover effects between treatment and control groups in a localized urban transportation network?

  • 02

    A VP wants a read in six weeks on a forward-stocking policy now live at eight of forty nodes. Replenishment lead time on the affected lanes is five weeks, the eight nodes were chosen by the ops team because they were the worst performers, and the cross-node standard deviation of weekly fill rate is about 3 points. The VP has already told his leadership team that a number is coming. Tell him what you can and cannot deliver in six weeks, and propose the design and timeline you would commit to instead.

  • 03

    Three requests land in the same week. Transport wants landed cost per delivered unit split by lane and mode from fct_shipment_leg for a carrier bid that closes in nine days. A plant manager wants OEE on two work centres decomposed into availability, performance and quality from fct_production_run for a capital request due next month. Commercial wants a fill-rate root cause from fct_order_line for an account review on Thursday. You are the only data scientist and have about four working days. Give your order, the rule that produced it, and what you say to the two who wait.

PracHub interview preparation framework ↗
Is this an official Waymo interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Waymo. Rounds and questions reflect what candidates have reported, not a process Waymo has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult is the interview loop at Waymo, and how much preparation time should I plan for?

The interview process is widely recognized as rigorous and challenging, particularly during the technical and statistical deep-dive rounds. Most successful candidates dedicate between four to six weeks of focused preparation, concentrating heavily on advanced statistics, experimental design pitfalls, and SQL window functions.

PracHub interview research ↗
What differentiates an average candidate from a top-tier candidate during the onsite loops?

Top candidates distinguish themselves by how they handle ambiguity. Instead of rushing into calculations, they pause to clarify assumptions, structure their approach logically, and connect their technical solutions back to the core safety and business objectives of Waymo. They also treat the interview as a collaborative engineering discussion.

PracHub interview research ↗
Are remote work options available for Data Scientists at Waymo?

Many Data Scientist roles operate on a hybrid schedule, typically based out of major hubs like Mountain View or San Francisco, California, though specific remote flexibilities vary by team and role level. Your recruiter will provide exact location and hybrid attendance guidelines during your initial screen.

PracHub interview research ↗
What is the typical timeline from an initial recruiter screen to a final offer decision?

The end-to-end timeline generally spans three to four weeks, moving from the recruiter chat and initial technical assessment through the technical phone screen and multi-round final onsite loop. However, scheduling adjustments can occasionally occur around holidays or interview panel availability.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.