At Saint-Gobain, a Data Scientist sits at the intersection of industrial tradition and digital transformation. As a world leader in sustainable construction and high-performance materials, the company relies on data to optimize complex manufacturing processes, reduce environmental impact, and streamline global supply chains. You aren't just building models in a vacuum; you are translating physical manufacturing challenges into mathematical solutions that affect real-world production lines and logistics networks.
The impact of this role is significant. Whether you are working on predictive maintenance for glass manufacturing equipment or optimizing the energy consumption of a production plant, your work directly contributes to Saint-Gobain’s goal of carbon neutrality. You will collaborate with multi-disciplinary teams, including engineers, plant managers, and product owners, to turn vast amounts of industrial data into actionable insights that drive strategic business decisions across the globe.
This position is ideal for those who enjoy the complexity of "noisy" real-world data and the challenge of deploying scalable solutions in a traditional industry. The scale of provides a unique playground where even a small percentage of optimization can lead to massive cost savings and a substantial reduction in the company's global carbon footprint.
Initial Screening
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Evaluation
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Problem-Solving Case Study
reportedThis round runs as a working session, so part of what it decides is whether you are useful to think with. The interviewer will interrupt: a hint that the data you assumed does not exist, a challenge to your metric, a nudge toward a branch you skipped. Treating those as interference is the common failure. Reason out loud while your thinking is still provisional so there is something to react to, and when a redirect arrives, take it instead of defending the path you had already started down.
What to demonstrate
- Whether your reasoning is audible while it is still unsettled, or only after you have privately decided
- What you do with a hint: absorb it and adjust, or argue for the original route
- Whether your clarifying questions have answers that would change your approach, as opposed to filling silence
- Whether you can be wrong about something in the middle of the case and keep moving without restarting
How to prepare
- Run practice cases with a partner instructed to interrupt twice: once to remove a data source you assumed existed, once to reject the metric you chose. Practise absorbing both without going back to the start.
- Before each practice case, write down the clarifying questions you plan to ask, then check afterwards whether any answer actually changed what you did. Drop the ones that did not.
- Explain an analysis you already know well to someone outside the field and have them stop you at every point where the reasoning jumped a step.
Panel Presentation
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
HR and Director Rounds
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
PracHub editorial advice for the preparation topics above.
Computing average inventory from period-end snapshots
Shipments cluster before period close, so the month-end on-hand position is systematically the lowest point of the month; turns computed against it are biased high and days of supply biased low, frequently by ten to twenty percent, and the bias grows precisely when close-period push is strongest. The same shape of error appears when a daily snapshot is joined to shipment events on date equality: a SKU with several legs on one day fans the snapshot out and multiplies the valued inventory. Average across every daily snapshot in the window for the denominator, and join snapshots to events with an explicit as-of condition and a row-count check at each grain before aggregating.
Sizing safety stock as z times sigma_D times the square root of lead time
That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.
Explaining an aggregate move without decomposing the mix shift
Split the change in the aggregate into within-segment movement and movement in segment weights before you explain it. Every segment's rate can fall while the overall rate rises, purely because volume shifted toward segments that already had higher rates.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Can you explain the bias-variance tradeoff?
Can you explain the bias-variance tradeoff?
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Translate the result into the decision it informs, in one plain sentence.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
What is the difference between bagging and boosting?
What is the difference between bagging and boosting?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
Describe a project where your initial model didn't perform as expected…
Describe a project where your initial model didn't perform as expected. How did you troubleshoot and iterate?
Approach
- Set a baseline first, so any model has something honest to beat.
- Check what information would not exist at prediction time, and exclude it.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Measure demand amplification hop by hop up the network
Using dim_location (location_id, parent_location_id, echelon, location_type) and fct_shipment_leg (leg_id, shipment_id, origin_location_id, destination_location_id, direction, tendered_at_utc, shipped_units), quantify how order variability grows upstream. fct_shipment_leg carries no sku_id and no link back to order lines, so shipped_units is the node's whole mixed-SKU flow and the measurement is node-level by construction. For each node compute weekly units flowing out (what it served downstream) and weekly units flowing in (what it ordered up), then the ratio of their variances on detrended, deseasonalised residuals. Walk parent_location_id from a store to echelon 0 and report the ratio at each hop, naming the hop where amplification is introduced. dim_location has no path column; build the chain yourself.
Approach
- Date orders by tendered_at_utc, not delivered_at_utc. Tender is the closest observable proxy for the moment the ordering decision was made; dating by delivery shifts the series by transit time and smears the variance you are trying to measure.
- Build both series on the same calendar weeks and in the same units. Variance scales with the aggregation window, so a ratio formed from daily inbound against weekly outbound measures the calendar, not the ordering rule. The units are mixed-SKU counts because the leg table carries no sku_id, so a node whose product mix drifts toward smaller or larger pack sizes moves both series for reasons unrelated to its ordering rule; print that caveat next to the number rather than implying a per-SKU result.
- Detrend and deseasonalise before taking variances, or compute the ratio on residuals from a simple weekly seasonal baseline. A growing node otherwise scores as amplifying, when all you have measured is its trend.
- Walk the parent chain iteratively: start from the store rows, join dim_location to itself on parent_location_id, and repeat until parent is null, capping the loop at the known echelon depth and asserting the path length matches the echelon difference so a data cycle raises rather than hangs.
- Read the output as a sequence, not a set of numbers. A pass-through node such as a cross-dock should sit near 1.0, and the hop where the ratio jumps is where a batching rule, a minimum order quantity or truckload rounding lives. That is the node to fix, not the node that is complaining.
Worked solution 45 min
- Assign an ISO week from tendered_at_utc and build two weekly series per node from fct_shipment_leg alone: outbound units summed where origin_location_id is the node, inbound units summed where destination_location_id is the node. No SKU filter is applied because the table has no sku_id, so both series are total unit flow.
- Regress each series on a linear trend plus week-of-year dummies, or subtract a centred moving average, and keep the residuals.
- Compute the amplification ratio per node as var(inbound residuals) / var(outbound residuals) over the shared weeks, requiring a minimum of 26 weeks before reporting a ratio.
- Build the parent chain from a chosen store by iteratively joining dim_location on parent_location_id until parent_location_id is null, asserting the loop terminates within the echelon depth.
- Emit the chain in order with echelon, location_type, both variances, the ratio, the week count and the mixed-SKU caveat, and mark the hop with the largest increase.
Follow-up
- The ratio at one hop is 4.2 and the node insists it orders exactly to forecast. What ordering rules would produce that number anyway?
- What would it take to run this for a single SKU family rather than total units, given fct_shipment_leg carries neither a sku_id nor any link to order lines, and which hops would still be unmeasurable after you added that link?
- What would you expect this measurement to look like during a promotion, and does that change your conclusion?
How would you handle a dataset that is too large to fit into memory?
How would you handle a dataset that is too large to fit into memory?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Explain the difference between a list and a tuple in Python and when y…
Explain the difference between a list and a tuple in Python and when you would use each.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
How would you join two large tables in SQL if you have a many-to-many …
How would you join two large tables in SQL if you have a many-to-many relationship?
Approach
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Write a function in Python to detect outliers in a list of sensor read…
Write a function in Python to detect outliers in a list of sensor readings.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Lag-7 forecast accuracy and bias at weekly grain
fct_inventory_daily has one row per sku_id x location_id x inventory_date carrying demand_qty, forecast_qty_lag7 (the forecast for that date as published seven days earlier, NULL before the series existed) and stockout_flag. Compute trailing eight-week WMAPE and signed weighted bias at the sku x location x week grain the ordering decision uses: aggregate demand and forecast to the week first, then form errors on the weekly totals. Exclude any week containing a NULL forecast day and report how many cell-weeks that removed. Return one accuracy row per ABC class.
Approach
- Roll daily rows to sku-location-week with SUM(demand_qty) and SUM(forecast_qty_lag7), carrying COUNT(*) FILTER (WHERE forecast_qty_lag7 IS NULL) so a partial week is visible rather than quietly summing to an artificially low forecast.
- Drop cell-weeks with any NULL forecast day and emit the dropped count as a column: an accuracy figure without its exclusion count cannot be checked by anyone.
- WMAPE = SUM(ABS(weekly_demand - weekly_forecast)) / NULLIF(SUM(weekly_demand), 0). Bias = SUM(weekly_forecast - weekly_demand) / NULLIF(SUM(weekly_demand), 0), signed, with cancellation as the whole point of reporting it next to WMAPE.
- Join dim_sku for abc_class and group there by re-summing both sides. Averaging per-SKU WMAPE gives a number weighted by nothing in particular.
- Report the share of retained cell-weeks containing a stockout day beside the result, because demand_qty on those days is censored at what could be supplied and the error is being scored against a truncated actual.
Worked solution 25 min
- CTE one: GROUP BY sku_id, location_id, DATE_TRUNC('week', inventory_date) with SUM(demand_qty), SUM(forecast_qty_lag7) and the NULL-day count.
- CTE two: split the cell-weeks into retained (null_days = 0) and excluded, and keep COUNT(*) of each.
- Join dim_sku on sku_id for abc_class, then aggregate the retained set per class with the two ratio formulas.
- Compute the stockout exposure share over the same retained set and attach it to each class row.
- Print retained count, excluded count and window bounds alongside every ratio.
Follow-up
- Why is MAPE unusable on a C-class item at a forward-stocking location, and what do you score it with instead?
- WMAPE improves sharply when you move from sku-location-week to sku-region-week. Which of the two belongs in the planning review, and why?
- Bias is +6 percent and WMAPE is 45 percent. What does that pair tell you to investigate first?
Size a fill-rate test when only nodes can be randomised
A new allocation rule can only be switched on per node. You have 24 regional DCs, 12 per arm, and eight post-period weeks. The outcome is weekly unit fill rate per node from fct_order_line: SUM(shipped_qty) over SUM(ordered_qty) by requested_ship_date week. Between-node SD of that weekly rate is 4.0 percentage points; within-node week-to-week SD is 1.5 points. Planning wants to detect a 1.5-point lift at 80% power, two-sided 5%. Give the minimum detectable effect for this design and say whether running longer fixes it.
Approach
- Write the variance at the unit that was actually randomised. The node is the unit, so each node contributes one number and the week rows inside it are not independent observations: the variance of one node's eight-week mean is sigma_b^2 + sigma_w^2/m, and the variance of an arm mean is that quantity over n nodes.
- Plug in: 4.0^2 + 1.5^2/8 = 16.28 pp^2. MDE = (1.96 + 0.84) * sqrt(2 * 16.28 / 12) = 2.80 * 1.647 = 4.6 pp.
- Re-run with m taken to infinity to show what duration buys: 2.80 * sqrt(2 * 16 / 12) = 4.57 pp. Between-node variance is 98.6% of the node-mean variance, and extra weeks only shrink the other 1.4%, so doubling the run changes the MDE by under 1%.
- Invert the formula for the nodes needed to reach 1.5 pp: 2 * 16.28 * (2.80/1.5)^2 = about 114 nodes per arm. That does not exist, so state plainly that the design cannot answer the question as posed rather than reporting an underpowered null.
- Offer the levers that attack sigma_b rather than time: matched-pair randomisation on pre-period fill rate with ANCOVA on the pre-period node rate (at a pre-post correlation of 0.8 the residual between-node variance is 5.76 and the MDE falls to about 2.8 pp), or a within-node switchback by week if the rule can be toggled.
Follow-up
- The rule can be toggled weekly. What does a node-level switchback buy you here, and what does it cost you in assumptions?
- If you must ship a decision on 24 nodes, what MDE do you pre-register, and what exactly does a null result license anyone to conclude?
Diagnose a sample ratio mismatch in an order-line test
A promise-date algorithm was randomised 50/50 with the hash taken on order_id. Your analysis table has 180,000 fct_order_line rows in the window: 91,540 control and 88,460 treatment. Lines are excluded when line_status = 'cancelled_by_customer' or when the join to fct_shipment_leg on shipment_id returns no delivered leg. The readout shows treatment up 0.9 points on on-time delivery. Test the split, state what the result implies about the readout, and list the three most likely causes given how this table was built.
Approach
- Run the chi-square goodness-of-fit test against the designed 50/50 split: expected 90,000 per arm, deviation 1,540, so chi-square = 2 * 1540^2 / 90000 = 52.7 on 1 degree of freedom, p around 4e-13. That is not sampling noise, and the readout is not interpretable until it is explained.
- Check the grain mismatch first, because it is the cheapest explanation: assignment is on order_id but rows are lines. A perfectly balanced order split still yields unequal line counts whenever lines per order differ by arm, and a promise-date change that splits or consolidates orders does exactly that. Recount distinct order_id per arm before touching lines.
- Walk the filter chain and recount the ratio at every step: raw assignment log, then all lines, then after the cancelled_by_customer filter, then after the delivered-leg join. The step where the ratio breaks names the cause without any further argument.
- Recognise that both exclusions are post-treatment. Requiring a delivered leg conditions on an outcome the treatment moves, which opens a collider path: whichever arm ships more of the difficult lines retains more slow lines and is penalised for succeeding.
- Do not report the 0.9-point lift. Rebuild the metric with all assigned orders in the denominator and never-shipped orders scored as failures, then re-run and compare.
Worked solution 20 min
- Compute the test: expected 90,000 per arm, deviation 1,540, chi-square = 2 * (1540^2 / 90000) = 52.7, p about 4e-13.
- Recount distinct order_id per arm on the raw assignment log, before any filter or join touches the data.
- Recount after each filter in the order the pipeline applies them, recording the arm ratio at each stage.
- If the raw order split is balanced, rebuild the readout at order grain with every assigned order in the denominator and unshipped orders counted as not on time.
Follow-up
- The split is clean at order_id but broken at line level. Is the readout salvageable, and at which grain would you report it?
- How would you monitor for this automatically on a test that runs for six weeks, and at what threshold would you halt?
Forecast accuracy collapsed at nodes missing snapshot rows
WMAPE at lag 7 for one region jumped from 31 percent to 58 percent over three weeks with no model change and no retraining. fct_inventory_daily is specified to carry one row per active sku-location-date even when on_hand_qty is zero, and a release changed how the end-of-day cut-off is derived from dim_location.timezone. Using fct_inventory_daily (inventory_date, sku_id, location_id, demand_qty, forecast_qty_lag7, stockout_flag) and dim_location, decide whether demand changed or the table did, and quantify the contaminated share before anyone retrains.
Approach
- Run a completeness check before touching any accuracy number: expected rows per location per day from the active sku-location grid, against actual rows present. If the table is incomplete the accuracy question is not well posed, and this check costs one query.
- Separate the two failure modes a cut-off change produces. Rows missing entirely shrink the denominator, and the sku-days that vanish are not a random sample. Rows present but with demand pushed onto the adjacent date create a one-day shift that inflates absolute error on both days while leaving the weekly total intact.
- Test the shift hypothesis directly by comparing SUM(demand_qty) per ISO week per node before and after the release. If weekly totals match and only the daily split moved, the model is fine and the grain is broken.
- Recompute WMAPE on clean node-days only and report both figures with the excluded row count and the share of demand those rows represent, since the metric definition already requires the excluded count to travel with the number.
- Check whether the affected nodes share a timezone that crosses a daylight-saving boundary or sits far from the offset the pipeline assumed. That grouping is the discriminating evidence between a code defect and genuine demand movement.
Follow-up
- Weekly totals match and only the daily split moved. Does the ordering decision care? At what replenishment frequency does it stop mattering?
- How would you backfill the missing days, and which of them would you refuse to backfill?
- What assertion would you add to the snapshot job so this fails loudly instead of surfacing three weeks later as model drift?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.
Tell me about your PhD research (or most significant project) and how …
Tell me about your PhD research (or most significant project) and how it can be applied to real-world problems.
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Quantify the outcome, including what you would not claim credit for.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Disagree with a planning manager about a safety stock cut
A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.
Approach
- Lead with the decision at risk, not the metric error: a 30 percent cut on intermittent items is the cheapest way to convert a measurement artefact into stockouts eight weeks later, by which time nobody will connect the two.
- Name the two defects precisely. Dropping zero-actual rows is not a rounding choice, it removes most of the history for a C-class item at a forward-stocking location and it removes the rows where over-forecasting is penalised, so the surviving MAPE is biased toward whichever series stocked out. Using shipped_qty as the actual scores the forecast against a supply ceiling, so a series looks more accurate the more often it ran out.
- Offer the replacement metric in the same breath: WMAPE, sum of absolute errors over sum of actuals against demand_qty, computed at the sku-location-week grain the ordering decision uses and at the lag it uses, reported next to signed bias so a persistently short series cannot hide inside an accuracy number.
- Propose the smaller action that is defensible now: recompute on the corrected basis, rank cells by bias rather than accuracy, and cut cover only where the corrected series is unbiased or over-forecasting, in a staged rollout with a service guardrail.
- Give him the win he actually wants: if the corrected numbers still support cuts on a subset, say so in advance, so the disagreement is about evidence rather than about territory.
Follow-up
- He says demand_qty is itself incomplete because customers stop ordering what shows out of stock. Is he right, and what do you do about it?
- Which items would you leave alone regardless of what the corrected metric says?
Defend a finding that the expedite program bought nothing
Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.
Approach
- Restate his objection in its strongest form before answering it, because a counterfactual worsening is a legitimate argument and treating it as an excuse ends the conversation.
- Separate what you measured from what you claimed: the data show no detectable service gain, not that expedite has no effect, and the difference is the entire argument.
- Test his hypothesis with the data you already have rather than defending in the abstract: split legs by exception_code = 'customs_hold' and by lane, and compare expedited against non-expedited legs on the same lanes in the same weeks, since if expedite were holding the line the expedited lanes should show a service gap over comparable non-expedited ones.
- Concede what is true: without a holdout you cannot rule out a protective effect, and the pre-period is contaminated by the disruption, so the honest statement is an upper bound on the gain rather than a zero.
- Hold the part that survives: the spend is real, it is concentrated in a small set of lanes, and nobody set a decision rule for when a leg gets expedited, which is a controllable problem independent of the counterfactual.
- Name the next measurement: a lane-level staggered switch-off with a stated burn-in, and say what it would cost and how long it would take.
Follow-up
- He offers to run the switch-off only on his two best lanes. Why is that a problem, and what do you counter with?
- His director asks you for a yes or no on cutting the budget today. What do you say?
- 01
Tell me about your PhD research (or most significant project) and how it can be applied to real-world problems.
- 02
A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.
- 03
Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.
Is this an official Saint-Gobain interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Saint-Gobain. Rounds and questions reflect what candidates have reported, not a process Saint-Gobain has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Data Scientist interview at Saint-Gobain?
The difficulty is generally rated as average to difficult. While the coding questions are often straightforward, the technical theory and the case study presentation require deep preparation and the ability to defend your logic under pressure.
PracHub interview research ↗What is the company culture like for the data team?
The culture is professional and collaborative. There is a strong emphasis on "competence" and "rigor." You will find that the team is very supportive, but they have high expectations for the quality of your work and your ability to deliver practical results.
PracHub interview research ↗How long does the entire interview process take?
The process typically takes between 3 to 6 weeks from the initial screen to the final offer. This depends on the location and the availability of the panel members for the presentation stage.
PracHub interview research ↗Is there a focus on specific tools?
Saint-Gobain uses a variety of tools, but Python, SQL, and Azure are very common. Being proficient in these will give you a significant advantage during the technical assessments.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22