As a Data Scientist at WestRock, you sit at the intersection of advanced analytics and industrial innovation. WestRock is a leader in sustainable paper and packaging solutions, and this role is pivotal in transforming massive datasets into actionable insights that optimize manufacturing processes, streamline supply chain logistics, and enhance customer experiences. Your work directly influences how the company innovates within a highly complex, global operational environment.
You will be responsible for building predictive models, developing analytical tools, and translating technical findings into strategic recommendations for stakeholders across the organization. Whether you are analyzing production efficiency or forecasting market trends, your contributions provide the empirical foundation for high-stakes business decisions. This role is designed for those who thrive on solving tangible, real-world problems and who are comfortable navigating the unique technical challenges of a large-scale industrial enterprise.
While the core of the role is technical, your ability to articulate the "why" behind your data models is just as important as the code itself. Be prepared to explain complex concepts to non-technical business partners.
Initial Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Assessments
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub editorial advice for the preparation topics above.
Sizing safety stock as z times sigma_D times the square root of lead time
That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.
Scoring intermittent demand with MAPE
MAPE is undefined whenever the actual is zero, which is most days for a C-class item at a forward-stocking location, and it is asymmetric even where it is defined: under-forecasting is bounded at 100 percent error while over-forecasting is unbounded. Optimising it therefore drives systematic under-forecasting on exactly the items whose stockouts are most expensive, and the bias is invisible in the headline because the zero-actual rows were dropped before averaging. Use WMAPE (sum of absolute errors over sum of actuals), a scaled error such as RMSSE, or a pinball loss at the service quantile the policy targets, and always state the aggregation level and forecast lag, because the same series scores very differently at daily SKU-store level and weekly SKU-region level.
Extrapolating a first-week lift inflated by novelty effects
Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
When faced with a messy or incomplete dataset, what is your standard a…
When faced with a messy or incomplete dataset, what is your standard approach to cleaning and feature engineering?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Say how the offline result would be validated online before it is trusted.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
Measure demand amplification hop by hop up the network
Using dim_location (location_id, parent_location_id, echelon, location_type) and fct_shipment_leg (leg_id, shipment_id, origin_location_id, destination_location_id, direction, tendered_at_utc, shipped_units), quantify how order variability grows upstream. fct_shipment_leg carries no sku_id and no link back to order lines, so shipped_units is the node's whole mixed-SKU flow and the measurement is node-level by construction. For each node compute weekly units flowing out (what it served downstream) and weekly units flowing in (what it ordered up), then the ratio of their variances on detrended, deseasonalised residuals. Walk parent_location_id from a store to echelon 0 and report the ratio at each hop, naming the hop where amplification is introduced. dim_location has no path column; build the chain yourself.
Approach
- Date orders by tendered_at_utc, not delivered_at_utc. Tender is the closest observable proxy for the moment the ordering decision was made; dating by delivery shifts the series by transit time and smears the variance you are trying to measure.
- Build both series on the same calendar weeks and in the same units. Variance scales with the aggregation window, so a ratio formed from daily inbound against weekly outbound measures the calendar, not the ordering rule. The units are mixed-SKU counts because the leg table carries no sku_id, so a node whose product mix drifts toward smaller or larger pack sizes moves both series for reasons unrelated to its ordering rule; print that caveat next to the number rather than implying a per-SKU result.
- Detrend and deseasonalise before taking variances, or compute the ratio on residuals from a simple weekly seasonal baseline. A growing node otherwise scores as amplifying, when all you have measured is its trend.
- Walk the parent chain iteratively: start from the store rows, join dim_location to itself on parent_location_id, and repeat until parent is null, capping the loop at the known echelon depth and asserting the path length matches the echelon difference so a data cycle raises rather than hangs.
- Read the output as a sequence, not a set of numbers. A pass-through node such as a cross-dock should sit near 1.0, and the hop where the ratio jumps is where a batching rule, a minimum order quantity or truckload rounding lives. That is the node to fix, not the node that is complaining.
Follow-up
- The ratio at one hop is 4.2 and the node insists it orders exactly to forecast. What ordering rules would produce that number anyway?
- What would it take to run this for a single SKU family rather than total units, given fct_shipment_leg carries neither a sku_id nor any link to order lines, and which hops would still be unmeasurable after you added that link?
- What would you expect this measurement to look like during a promotion, and does that change your conclusion?
Cluster bootstrap for landed cost per delivered unit
You have delivered outbound legs from fct_shipment_leg: leg_id, shipment_id, leg_seq, carrier_id, origin_location_id, destination_location_id, shipped_units, cube_m3, linehaul_cost_cents, fuel_surcharge_cents, accessorial_cost_cents, expedite_premium_cents. In this extract the linehaul for a multi-stop load is booked entirely on leg_seq 1, so allocate it across the load's legs by cube first. Then estimate the difference in landed cost per delivered unit between two carriers on lanes both serve, with a 95 percent interval. Write the resampling yourself; no bootstrap library.
Approach
- Allocate the linehaul before anything else, in proportion to each leg's cube over the load's total cube, and keep the other three cost columns where they already sit. State the choice: allocating by weight instead reorders lanes whenever freight is bulky rather than dense.
- Restrict to the set of lanes both carriers actually serve in the window. Comparing over all lanes measures which carrier was assigned the cheap lanes, not which carrier is cheaper.
- Resample shipments, not legs. Legs of one load share a single allocated linehaul and a single dispatch decision, so treating them as independent draws understates the variance of the estimate.
- Recompute the statistic as a ratio of sums on every resample, total allocated cost over total shipped units. Taking the mean of leg-level cost-per-unit values instead gives a different estimand in which a one-unit leg counts as much as a full truckload.
- Report both the raw difference and a lane-standardised difference where lane weights are fixed at the pooled volume share, and say which one you would put in front of a decision maker.
Worked solution 35 min
- Compute per-shipment total cube, allocate the leg_seq 1 linehaul across legs by cube share, and form leg_total_cost from the four cost columns.
- Build the lane key from origin_location_id and destination_location_id, and keep only lanes with delivered volume from both carriers.
- Compute the point estimate per carrier as sum(leg_total_cost) / sum(shipped_units), and take the difference.
- Draw 2,000 bootstrap replicates by sampling shipment_ids with replacement within each carrier, rebuilding both sums from the sampled legs and recomputing the ratio difference.
- Take the 2.5th and 97.5th percentiles of the replicate differences, then repeat the whole procedure with resampling stratified inside lane to produce the mix-standardised interval.
Follow-up
- The interval crosses zero. What would you need in volume or in window length to resolve a difference of 2 cents per unit?
- One carrier's accessorials are rising while its linehaul is flat. What is that signature telling you about execution versus rates?
- How does your interval change if one shipment accounts for 15 percent of the units?
Perfect order rate against the original promise in local time
Compute weekly perfect order rate at order grain. An order qualifies only if every fct_order_line has line_status = 'shipped_full', every line's actual_delivery_at_utc falls on or before original_promised_delivery_date at the ship-from node's local end of day (dim_location.timezone holds IANA names), and no fct_shipment_leg reached through shipment_id carries exception_code in ('damage','address_issue') or leg_status in ('refused','lost'). The denominator is orders whose original_promised_delivery_date lands in the week. Undelivered lines fail. Publish with a fourteen-day lag and declare invoice accuracy as out of scope.
Approach
- Convert the event into local time rather than comparing against a UTC-anchored deadline: (actual_delivery_at_utc AT TIME ZONE 'UTC' AT TIME ZONE loc.timezone)::date <= original_promised_delivery_date. Casting the promise date to a timestamp puts the cut-off at UTC midnight starting the promise day, which is up to a full day early and wrong in both directions depending on the node's offset.
- Evaluate the three components as booleans at line grain, then roll to the order with BOOL_AND, coercing a NULL actual_delivery_at_utc to FALSE explicitly so an undelivered line fails rather than propagating NULL and dropping the order out of the comparison.
- Test the shipment exceptions with EXISTS against fct_shipment_leg, never with a join. One shipment has many legs and one order has many lines, so a join multiplies the order into the denominator and changes the metric.
- Aggregate to order grain first, then count qualifying orders over orders promised in the week. The rate is orders over orders; a line-level average is a different and easier metric.
- Apply the publication lag by restricting to weeks whose promise dates are at least fourteen days old, and emit the three component pass rates beside the headline so any movement is attributable to completeness, timeliness or exceptions.
Follow-up
- All three component rates are flat but perfect order fell two points. How is that arithmetically possible, and what does it tell you about where the failures now land?
- A node sits at UTC+13 and many orders are promised on the last day of the month. What does that do to your weekly and monthly buckets?
- How do you handle an order whose lines ship from two nodes in different timezones, and which node's clock owns the promise?
Stockout streaks of three days or more per cell
fct_inventory_daily carries one row per sku_id x location_id x inventory_date for every active pair, including days with nothing on hand, plus stockout_flag marking days where available_qty reached zero. Return every run of consecutive stockout days lasting three days or more: sku, location, start date, end date, run length, and demand_qty summed across the run. Return alongside it, per sku-location, the count of such runs in the window. Verify that the daily row series is contiguous for each pair before relying on date arithmetic, and say what you do about pairs that fail.
Approach
- Filter to stockout_flag = TRUE, then assign ROW_NUMBER() OVER (PARTITION BY sku_id, location_id ORDER BY inventory_date).
- Form the island key as inventory_date - rn * INTERVAL '1 day'. Two rows share a key only when their dates are exactly consecutive, so the key cannot weld separate runs together; the exposure runs the other way. A day with no snapshot row inside a true stockout breaks that run into fragments, and any fragment under three days is deleted by the threshold, so the run leaves the output rather than appearing with a wrong length.
- Prove contiguity before trusting the key: per sku-location compare COUNT(DISTINCT inventory_date) against (MAX(inventory_date) - MIN(inventory_date) + 1), and separately confirm no pair carries two rows on one date, since a duplicate date advances rn without advancing the calendar and shifts every key after it. Pairs that fail get a generated date spine: LEFT JOIN generate_series over the window so each day is labelled stockout, in stock, or no snapshot, then apply a stated rule for the no-snapshot days instead of letting the key make that call silently. The conservative rule is to break the run at the hole and mark both fragments gap_adjacent.
- Group by sku, location and island key for MIN and MAX date, COUNT(*) as length and SUM(demand_qty) as the demand exposed to the stockout, keeping runs of length >= 3.
- Aggregate once more per sku-location for the run count, and mark runs touching either window edge as censored, since their true length is unknown from this window alone.
Worked solution 30 min
- CTE completeness: per sku_id, location_id compute COUNT(DISTINCT inventory_date) and the date span, and flag pairs where they disagree or where COUNT(*) exceeds COUNT(DISTINCT inventory_date).
- CTE flagged: SELECT the stockout rows with ROW_NUMBER partitioned by sku_id, location_id ordered by inventory_date.
- CTE islands: add grp = inventory_date - rn * INTERVAL '1 day'.
- GROUP BY sku_id, location_id, grp producing MIN, MAX, COUNT() and SUM(demand_qty); filter COUNT() >= 3.
- Second aggregate over that result for runs per sku-location, plus a censored flag where start equals the window start or end equals the window end.
Follow-up
- The longest runs cluster on X-class A items. What does that combination suggest to look at before touching the safety stock policy?
- How would you estimate lost demand across a run, given that demand_qty records the orders that were placed but not the orders nobody bothered to place?
- The snapshot is taken at each location's local end of day and the pairs span several timezones. What breaks, and what stays true?
How have you utilized [specific programming language] to solve a real-…
How have you utilized [specific programming language] to solve a real-world business problem in your previous roles?
Approach
- Say what you would check first and why it is the highest-information step.
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
What specific software packages and libraries are you most proficient …
What specific software packages and libraries are you most proficient in for data manipulation?
Approach
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
How do you validate the performance of your models before deploying th…
How do you validate the performance of your models before deploying them into a production environment?
Approach
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Measure the demand that never becomes an order line
Commercial asks for lost sales by SKU and node. fct_order_line records short-shipped lines (short_reason_code = 'no_stock') and no-supply cancellations, and fct_inventory_daily carries demand_qty, available_qty and stockout_flag, but an order a customer never placed because the item showed unavailable leaves no row anywhere. Define the metric you will publish, the proxy it rests on, the direction and likely size of its bias, and the guardrail that stops the proxy being quoted as a measurement.
Approach
- Separate observable from unobservable in writing. Short-shipped units with short_reason_code = 'no_stock' and cancelled_qty on lines with line_status = 'cancelled_no_supply' are recorded demand that went unmet. Demand that never became a line is absent from the schema at every grain, so no query returns it and no join recovers it.
- Publish the observable part under its own name, unmet recorded demand, and refuse to call it lost sales, because the two quantities differ by exactly the amount nobody can see.
- Build the proxy from uncensored periods: per sku-location, estimate a baseline demand rate from days where stockout_flag is false and available_qty stayed comfortably above that day's demand, adjust for day of week and season, impute the censored days, and take the excess over recorded demand.
- State the bias and its direction. The estimate is a lower bound, because purchases diverted before an order was keyed and demand absorbed by substitution onto another SKU are both invisible. The shortfall is largest on the SKUs with the most stockout days, so it correlates with the quantity being measured and cannot be treated as noise that averages out.
- Bound the size instead of asserting it: report the censored share of sku-location-days, and run the imputation on a holdout of uncensored days that you censor artificially, so the recovery rate is measured on data where the truth is known.
- Put the guardrail on use rather than on the number: any spend justified by this metric must still hold when restated on the observable part alone, otherwise the proxy is carrying the argument and should be labelled as doing so.
Worked solution 40 min
- Quantify the censoring: share of sku-location-days with stockout_flag true, and share of ordered units on lines with short_reason_code = 'no_stock' or line_status = 'cancelled_no_supply'.
- Compute the observable metric: unmet recorded demand in units and at unit_price_cents, weekly by node and ABC class.
- Fit the baseline rate on uncensored days only, with day-of-week and week-of-year terms, pooled within ABC/XYZ class where a sku-location series is too thin to fit alone.
- Impute the censored days, take the gap, then hold out a random sample of uncensored days, censor them artificially at plausible availability levels, and measure what fraction of known demand the method recovers.
- Publish both series with the censored share and one written sentence naming what the proxy cannot see.
Follow-up
- substituted_sku_id records substitutions. How much of the gap does it close, and what does counting it break in the donor SKU's own demand history?
- How do you use the imputed series in a forecast without letting the imputation feed its own next round?
Forecast accuracy collapsed at nodes missing snapshot rows
WMAPE at lag 7 for one region jumped from 31 percent to 58 percent over three weeks with no model change and no retraining. fct_inventory_daily is specified to carry one row per active sku-location-date even when on_hand_qty is zero, and a release changed how the end-of-day cut-off is derived from dim_location.timezone. Using fct_inventory_daily (inventory_date, sku_id, location_id, demand_qty, forecast_qty_lag7, stockout_flag) and dim_location, decide whether demand changed or the table did, and quantify the contaminated share before anyone retrains.
Approach
- Run a completeness check before touching any accuracy number: expected rows per location per day from the active sku-location grid, against actual rows present. If the table is incomplete the accuracy question is not well posed, and this check costs one query.
- Separate the two failure modes a cut-off change produces. Rows missing entirely shrink the denominator, and the sku-days that vanish are not a random sample. Rows present but with demand pushed onto the adjacent date create a one-day shift that inflates absolute error on both days while leaving the weekly total intact.
- Test the shift hypothesis directly by comparing SUM(demand_qty) per ISO week per node before and after the release. If weekly totals match and only the daily split moved, the model is fine and the grain is broken.
- Recompute WMAPE on clean node-days only and report both figures with the excluded row count and the share of demand those rows represent, since the metric definition already requires the excluded count to travel with the number.
- Check whether the affected nodes share a timezone that crosses a daylight-saving boundary or sits far from the offset the pipeline assumed. That grouping is the discriminating evidence between a code defect and genuine demand movement.
Follow-up
- Weekly totals match and only the daily split moved. Does the ordering decision care? At what replenishment frequency does it stop mattering?
- How would you backfill the missing days, and which of them would you refuse to backfill?
- What assertion would you add to the snapshot job so this fails loudly instead of surfacing three weeks later as model drift?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.
Explain forecast uncertainty to a non-technical general manager
A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.
Approach
- Convert the error band into the units the approver already thinks in: weeks of cover, pallets, or dollars at risk, not a percentage or a confidence interval.
- Answer the right-or-wrong question directly rather than deflecting: the forecast will be wrong, the size of the wrongness is what you measured, and the build is sized against that size.
- State precisely what the 97 percent buys: it is a cycle service level, the probability of not running out during one replenishment cycle, so roughly three cycles in a hundred see a stockout. It is not the share of units shipped from stock. Unit fill rate also depends on the replenishment quantity, and when that quantity is large relative to the standard deviation of lead-time demand, which is the normal case outside lot-for-lot ordering, the fill rate sits above the cycle service number, often above 99 percent at a 97 percent cycle target.
- Give the two-sided consequence in money: what the build costs to carry at the applicable cost of capital plus obsolescence risk on shelf-life items, against the margin at risk from the stockouts it prevents.
- Close with the decision you want and the trigger that would reverse it, for example a lag-7 bias check after four weeks that reopens the number.
Follow-up
- The GM says just give me one number. What do you give, and what do you refuse to give?
- How would your answer change if the items were frozen with a 90-day shelf life?
Defend a finding that the expedite program bought nothing
Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.
Approach
- Restate his objection in its strongest form before answering it, because a counterfactual worsening is a legitimate argument and treating it as an excuse ends the conversation.
- Separate what you measured from what you claimed: the data show no detectable service gain, not that expedite has no effect, and the difference is the entire argument.
- Test his hypothesis with the data you already have rather than defending in the abstract: split legs by exception_code = 'customs_hold' and by lane, and compare expedited against non-expedited legs on the same lanes in the same weeks, since if expedite were holding the line the expedited lanes should show a service gap over comparable non-expedited ones.
- Concede what is true: without a holdout you cannot rule out a protective effect, and the pre-period is contaminated by the disruption, so the honest statement is an upper bound on the gain rather than a zero.
- Hold the part that survives: the spend is real, it is concentrated in a small set of lanes, and nobody set a decision rule for when a leg gets expedited, which is a controllable problem independent of the counterfactual.
- Name the next measurement: a lane-level staggered switch-off with a stated burn-in, and say what it would cost and how long it would take.
Follow-up
- He offers to run the switch-off only on his two best lanes. Why is that a problem, and what do you counter with?
- His director asks you for a yes or no on cutting the budget today. What do you say?
Scope a vague request about rising inventory
A planning director opens with: inventory is up twelve percent on flat shipments, get me an analysis by Friday. You have fct_inventory_daily (inventory_date, sku_id, location_id, on_hand_qty, standard_cost_cents) and fct_order_line (shipped_qty, unit_cogs_cents, requested_ship_date). Nothing else is specified: not the window, not the comparison basis, not whether the twelve percent is units or value. State the three questions you would ask before writing any SQL, name the decision each answer changes, and describe the first cut you would run if nobody answers you before Friday.
Approach
- Ask what decision hangs on the answer: a buy stop, a write-off provision, a policy review and a board slide all need different cuts, and the director usually has one of them in mind.
- Pin the measurement before the cause: units or value, which two windows are being compared, and whether the denominator is average daily on-hand or a period-end snapshot, since a month-end denominator biases turns high by ten to twenty percent and can manufacture the whole movement.
- Ask which part of the network is in scope, because echelon matters: a build at echelon 2 ahead of a promotion and a build of blocked_qty at a plant are different problems with different owners.
- Commit to a default if no answer arrives: value at standard cost, averaged across every daily snapshot in both windows, cut by echelon, then by abc_class and xyz_class, then by lifecycle_status to separate phase_out stock from active cover.
- Say out loud what the first cut cannot settle, so the director is not surprised: it localises the build but does not attribute it to forecast bias, a lot-size change or a supplier pulling orders in.
Follow-up
- The build is concentrated in one product family at two DCs. What are your next two queries, and what would make you stop calling it a planning problem?
- The director wants the number by Friday and the cause by Friday. Which do you drop, and how do you say so?
- 01
A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.
- 02
Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.
- 03
A planning director opens with: inventory is up twelve percent on flat shipments, get me an analysis by Friday. You have fct_inventory_daily (inventory_date, sku_id, location_id, on_hand_qty, standard_cost_cents) and fct_order_line (shipped_qty, unit_cogs_cents, requested_ship_date). Nothing else is specified: not the window, not the comparison basis, not whether the twelve percent is units or value. State the three questions you would ask before writing any SQL, name the decision each answer changes, and describe the first cut you would run if nobody answers you before Friday.
Is this an official WestRock interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at WestRock. Rounds and questions reflect what candidates have reported, not a process WestRock has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
While it varies by team and location, the process is generally completed within a few weeks across three core stages. Stay in close contact with your recruiter for updates on your specific timeline.
PracHub interview research ↗Is the technical interview focused on whiteboard coding or project discussion?
Expect a hybrid approach. You will likely be asked to discuss your past projects in detail, but you should also be prepared for technical questions regarding software packages and data manipulation methodologies.
PracHub interview research ↗What is the best way to stand out?
The most successful candidates are those who can clearly link their technical skills to business outcomes. Focus on the impact your projects had on the company’s bottom line or operational efficiency.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22