XPO · Data Scientist
Updated · 2026-09-22

XPO Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at XPO, you will sit at the intersection of cutting-edge technology and global supply chain logistics. XPO is one of the largest providers of asset-based less-than-truckload (LTL) transportation in North America. In this role, your work directly impacts how millions of tons of freight move across the continent daily. You will build and deploy models that solve highly complex, real-world physical challenges, such as route optimization, dynamic pricing, freight capacity forecasting, and labor scheduling.

A large share of questions open as "how would you measure X", where the real work is choosing the metric, fixing its denominator, and defining the population it applies to. Any computation comes last and is frequently not required at all.

XPO candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Size safety stock under variable lead timesJoin daily snapshots to shipment events safelyPrice a service-level change in working capital

33 min read

Practice 13 Data Scientist prompts
13Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at XPO, you will sit at the intersection of cutting-edge technology and global supply chain logistics. XPO is one of the largest providers of asset-based less-than-truckload (LTL) transportation in North America. In this role, your work directly impacts how millions of tons of freight move across the continent daily. You will build and deploy models that solve highly complex, real-world physical challenges, such as route optimization, dynamic pricing, freight capacity forecasting, and labor scheduling.

The data ecosystem at XPO is massive, fast-moving, and highly complex. Unlike digital-only products, the variables you model here have physical consequences—ranging from fuel consumption and transit times to warehouse space utilization and carrier efficiency. This means your predictive models and optimization algorithms must be robust, scalable, and highly accurate, as they directly influence operational decisions and bottom-line profitability.

For a data professional, this environment offers an incredibly rich playground. You will collaborate closely with operations, engineering, and product teams to translate physical logistics problems into mathematical frameworks. If you are excited by the prospect of seeing your algorithms optimize massive fleets of trucks and streamline multi-million-dollar supply chain operations, the role at offers an unparalleled opportunity for high-impact work.

01

HR Recruiter Screen

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research
02

Initial Technical Screening

reported

Much of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.

What to demonstrate

  • Whether the query you write matches the plan you just described
  • What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
  • Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly

How to prepare

  • Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
  • Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
  • Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
PracHub interview research
03

Core Interview Rounds

reported

An extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.

What to demonstrate

  • Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
  • Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
  • How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one

How to prepare

  • Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
  • Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
  • Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
PracHub interview research
04

Take-Home Data Project

reported

A take-home is graded as an argument, not as a notebook. Somebody reads the submission without you in the room, so every choice has to survive on the page: why the question was framed this way, and what was deliberately left out. The gap between a strong and a weak submission is almost never model quality. It is whether the writeup names the specific question it answers and commits to a recommendation, including what evidence would overturn it. A high-accuracy model attached to no conclusion reads as effort that stopped before the decision.

What to demonstrate

  • Whether the question you answered is stated outright, and whether it is the question the prompt posed rather than an easier neighbour of it
  • Whether the recommendation is specific enough to act on, with the uncertainty attached to it instead of parked in a caveats section at the end
  • Whether analytical choices such as the metric definition, the population filter and the time window are justified in the prose, not merely visible in code

How to prepare

  • Take a dataset you have already worked with, write the one-paragraph conclusion first, then check whether the analysis you were planning actually supports it and cut whatever does not
  • Practise stating a metric in one sentence that fixes the population, the time window and the denominator, then confirm your query computes exactly that sentence and nothing adjacent to it
  • Hand a draft to someone outside the problem and ask them to tell you back what you recommended and why; anything they cannot recover is not on the page yet
PracHub interview research
05

Business-Focused Discussion

reported

An extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.

What to demonstrate

  • Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
  • Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
  • How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one

How to prepare

  • Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
  • Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
  • Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
PracHub interview research

PracHub editorial advice for the preparation topics above.

01

Attributing demand variability to the node where it is observed

Order variance amplifies as it moves upstream: batching to a truckload, minimum order quantities, forecast-driven ordering and promotional pull all convert smooth end demand into lumpy replenishment orders, so a plant can see a coefficient of variation several times that of the underlying consumption. Diagnosing the plant's schedule instability as a plant problem then produces interventions that cannot work, because the generating process sits one or two echelons downstream. Measure the bullwhip ratio explicitly (variance of orders placed by a node over variance of demand it received) at each echelon, and fix the ordering rule at the node where the ratio jumps rather than the node where the pain is felt.

02

Computing average inventory from period-end snapshots

Shipments cluster before period close, so the month-end on-hand position is systematically the lowest point of the month; turns computed against it are biased high and days of supply biased low, frequently by ten to twenty percent, and the bias grows precisely when close-period push is strongest. The same shape of error appears when a daily snapshot is joined to shipment events on date equality: a SKU with several legs on one day fans the snapshot out and multiplies the valued inventory. Average across every daily snapshot in the window for the denominator, and join snapshots to events with an explicit as-of condition and a row-count check at each grain before aggregating.

03

Reading experiment results before checking the arm split

Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.

04

Naming a model class before naming the deployment constraints

Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

10 technical prompts3 include a worked solution

How does a random forest algorithm differ from gradient boosting, and …

medium
machine learning and modelling

How does a random forest algorithm differ from gradient boosting, and in what scenarios would you prefer one over the other?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Say how the offline result would be validated online before it is trusted.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

Imagine we want to build a model to predict whether a shipment will be…

medium
machine learning and modelling

Imagine we want to build a model to predict whether a shipment will be delayed. What features would you engineer, and how would you handle highly imbalanced data?

Approach
  1. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  2. Pick an evaluation metric that matches the cost of each error type, not a default.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

Reconstruct replenishment cycles from a daily inventory series

mediumWorked solution
event sessionisationstate reconstructionlead time

For one sku-location you have a daily series: inventory_date, on_hand_qty, in_transit_qty, on_order_qty, allocated_qty, shipped_qty, demand_qty, reorder_point_qty, stockout_flag. There is no purchase-order table. Infer the replenishment cycles: a cycle opens on the first day inventory position falls to or below reorder_point_qty, and closes on the first subsequent receipt. Receipts must be inferred from the series itself. Return one row per cycle with trigger date, receipt date, inferred lead time in days, demand during the cycle and stockout days, then report mean and standard deviation of the inferred lead time.

Approach
  1. Derive the receipt signal from the stock balance rather than looking for a column that does not exist: on_hand today equals on_hand yesterday plus receipts minus units shipped, so implied_receipt = on_hand_qty.diff() + shipped_qty, and a positive value means goods arrived.
  2. State what else can make that quantity positive: a cycle-count adjustment, a customer return, or a scrap reversal. Corroborate each candidate receipt with a matching fall in in_transit_qty on the same or adjacent day, and discard the ones with no corroboration.
  3. Debounce arrivals. A load split across two dock shifts produces positive implied receipts on consecutive days; collapse runs of receipt days separated by less than a chosen gap into a single arrival and make the gap an explicit parameter.
  4. Compute inventory position as on_hand + in_transit + on_order - allocated, not on hand alone, or every day after a trigger looks like another trigger and the cycle count explodes.
  5. Walk the series once with an explicit state variable (awaiting-trigger, awaiting-receipt) rather than trying to vectorise the pairing. The state machine is what makes cycles non-overlapping, and it is what a reviewer can check line by line. Be clear that cycles defined this way do not tile the calendar: the stretch from a receipt until the position next falls to the reorder point sits inside no cycle.
Worked solution 35 min
  1. Sort by inventory_date, assert no duplicate dates and no calendar gaps, and compute position = on_hand_qty + in_transit_qty + on_order_qty - allocated_qty.
  2. Compute implied_receipt = on_hand_qty.diff() + shipped_qty and mark receipt days where it is positive and in_transit_qty fell.
  3. Collapse consecutive receipt days into single arrivals using a gap threshold of one day.
  4. Iterate the series with a two-state machine, emitting a cycle row when a trigger is followed by an arrival, and carrying an open trigger forward across days.
  5. For each cycle sum demand_qty and stockout_flag between trigger and receipt, then compute mean and standard deviation of the lead times across cycles.
EXPECTED RESULTA cycle table with non-overlapping intervals ordered in time, each with receipt_date strictly after trigger_date, plus two scalars: mean inferred lead time and its standard deviation across cycles. The cycles cover only the trigger-to-receipt stretches, so they will not add up to the length of the series; the post-receipt stretches where stock is still above the reorder point are outside every cycle by definition.
Follow-up
  • Two orders are open at once because the trigger fired again before the first arrived. How does your state machine change?
  • The policy quotes lead time in working days. What do you need from dim_location to convert, and where does the timezone matter?
  • You now have mean and standard deviation of lead time. Write the safety stock this implies and say which assumption you are least comfortable with.

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Data people depend on systems owned by other teams, and much of the job is negotiating for instrumentation, access, or a fix to a broken pipeline. Prepare an example of getting something changed upstream that you did not control. Describe what you asked for, what you traded, and how you worked while you waited.

How do you handle missing or corrupted GPS coordinates in a shipment t…

medium
behavioural and stakeholder questions

How do you handle missing or corrupted GPS coordinates in a shipment tracking dataset using Python?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Pick a story where you drove the decision, not one where you observed it.
  3. Close with what you would do differently, concretely.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

Describe a time when you had to present a complex machine learning mod…

medium
behavioural and stakeholder questions

Describe a time when you had to present a complex machine learning model to a non-technical stakeholder. How did you structure your explanation?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Close with what you would do differently, concretely.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Quantify your own impact on a safety stock change honestly

hard
impact measurementconfoundingdifference in differences

In March you re-sized safety stock at twelve regional DCs using the variable lead time formula. Unit fill rate at those nodes rose from 94.1 to 96.5 percent by June. In the same window a supplier that had been running 71 percent inbound on-time recovered to 93 percent, and Q2 volume is seasonally about 8 percent below Q1. Your manager asks for one sentence on impact for a performance review. Write the claim you are willing to defend under challenge, and explain how you bounded the portion attributable to your change.

Approach
  1. Refuse the naive claim first: the 2.4 point gross movement contains at least three generating processes, and attributing all of it is the sort of claim that collapses the moment someone plots the supplier's on-time series next to it.
  2. Build a comparison group from nodes that did not get the change, check that their pre-period fill rate moved in parallel with the treated nodes for at least two quarters before March, and estimate the difference in differences with variance clustered at the node, since twelve treated clusters is the effective sample size regardless of how many node-weeks there are.
  3. Strip the confounders you can measure directly rather than trusting the design alone: restrict to sku-location cells sourced from other suppliers to remove the recovery, and hold the sku mix fixed across periods so the seasonal volume drop does not enter through mix.
  4. Sanity-check the mechanism, not just the coefficient: if your change worked, the improvement should concentrate in cells whose reorder point rose and in short_reason_code = 'no_stock' lines, and should be absent where the policy did not move. If the gain is spread evenly, something else caused it.
  5. State the result as an interval with the confounders named, and state the cost side too: extra units held at the applicable cost of capital, so the claim is a net one rather than a service headline.
  6. Write the sentence so that it survives someone else re-running it, which usually means a smaller number than the gross movement and a named method.
Follow-up
  • The comparison nodes are not parallel in the pre-period. What do you do next?
  • Your bounded estimate is roughly half the gross movement. How do you present that to a manager who already told his boss the larger number?
  • 01

    How do you handle missing or corrupted GPS coordinates in a shipment tracking dataset using Python?

  • 02

    Describe a time when you had to present a complex machine learning model to a non-technical stakeholder. How did you structure your explanation?

  • 03

    In March you re-sized safety stock at twelve regional DCs using the variable lead time formula. Unit fill rate at those nodes rose from 94.1 to 96.5 percent by June. In the same window a supplier that had been running 71 percent inbound on-time recovered to 93 percent, and Q2 volume is seasonally about 8 percent below Q1. Your manager asks for one sentence on impact for a performance review. Write the claim you are willing to defend under challenge, and explain how you bounded the portion attributable to your change.

PracHub interview preparation framework
Is this an official XPO interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at XPO. Rounds and questions reflect what candidates have reported, not a process XPO has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research
How technical is the initial coding exam?

A: The initial exam is highly focused on speed and foundational accuracy. You will have 30 minutes for general aptitude (logical reasoning and math) and 30 minutes for multiple-choice questions on SQL and Python concepts. It is not an intense LeetCode-style algorithms test, but you must know your syntax and basic data operations thoroughly to finish within the time limit.

PracHub interview research
What is the work culture like for the Data Science team at XPO?

A: The culture is highly pragmatic, fast-paced, and collaborative. Because XPO is a physical logistics company, there is a strong emphasis on practical execution and measurable business impact over theoretical perfection. Teams work closely with operations, meaning you will get to see the physical results of your models in action very quickly.

PracHub interview research
How much domain knowledge in logistics do I need to have before interviewing?

A: While prior supply chain or logistics experience is a significant plus, it is not strictly required. However, you should show a genuine curiosity about how logistics networks operate. Showing that you understand concepts like hub-and-spoke networks, capacity constraints, and routing challenges during your case study discussions will make a very positive impression on the hiring team.

PracHub interview research
What is the typical timeline from the first screen to an offer?

A: The entire process generally takes between three to five weeks. This timeline can vary depending on whether the team requires a take-home project presentation and how quickly interviews can be scheduled with key business stakeholders. Recruiter communication is typically highly supportive and transparent throughout the process.

PracHub interview research
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.