WestRock · Data Scientist
Updated · 2026-09-24

WestRock Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at WestRock, you sit at the intersection of advanced analytics and industrial innovation. WestRock is a leader in sustainable paper and packaging solutions, and this role is pivotal in transforming massive datasets into actionable insights that optimize manufacturing processes, streamline supply chain logistics, and enhance customer experiences. Your work directly influences how the company innovates within a highly complex, global operational environment.

Allocate prep to your weakest link rather than your favourite topic. Of the three things that usually gate the outcome (SQL that is correct under messy joins, sound reasoning about experiments, and structured framing of an open-ended problem), candidates tend to over-invest in modelling theory and under-invest in framing.

WestRock candidates report 2 rounds · ≈ 2-4 weeks. The stages below are what candidates describe, not a published process.

Price a service-level change in working capitalSize safety stock under variable lead timesTrace an on-time miss to one node

35 min read

Practice 13 Data Scientist prompts
13Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at WestRock, you sit at the intersection of advanced analytics and industrial innovation. WestRock is a leader in sustainable paper and packaging solutions, and this role is pivotal in transforming massive datasets into actionable insights that optimize manufacturing processes, streamline supply chain logistics, and enhance customer experiences. Your work directly influences how the company innovates within a highly complex, global operational environment.

You will be responsible for building predictive models, developing analytical tools, and translating technical findings into strategic recommendations for stakeholders across the organization. Whether you are analyzing production efficiency or forecasting market trends, your contributions provide the empirical foundation for high-stakes business decisions. This role is designed for those who thrive on solving tangible, real-world problems and who are comfortable navigating the unique technical challenges of a large-scale industrial enterprise.

While the core of the role is technical, your ability to articulate the "why" behind your data models is just as important as the code itself. Be prepared to explain complex concepts to non-technical business partners.

01

Initial Screening

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Assessments

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Sizing safety stock as z times sigma_D times the square root of lead time

That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.

02

Scoring intermittent demand with MAPE

MAPE is undefined whenever the actual is zero, which is most days for a C-class item at a forward-stocking location, and it is asymmetric even where it is defined: under-forecasting is bounded at 100 percent error while over-forecasting is unbounded. Optimising it therefore drives systematic under-forecasting on exactly the items whose stockouts are most expensive, and the bias is invisible in the headline because the zero-actual rows were dropped before averaging. Use WMAPE (sum of absolute errors over sum of actuals), a scaled error such as RMSSE, or a pinball loss at the service quantile the policy targets, and always state the aggregation level and forecast lag, because the same series scores very differently at daily SKU-store level and weekly SKU-region level.

03

Extrapolating a first-week lift inflated by novelty effects

Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.

04

Reporting a p-value with no effect size or interval

Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

10 technical prompts3 include a worked solution

When faced with a messy or incomplete dataset, what is your standard a…

medium
machine learning and modelling

When faced with a messy or incomplete dataset, what is your standard approach to cleaning and feature engineering?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Say how the offline result would be validated online before it is trusted.
  3. Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • What would you monitor after launch to know the model is still valid?

Measure demand amplification hop by hop up the network

hard
bullwhiphierarchy traversalvariance

Using dim_location (location_id, parent_location_id, echelon, location_type) and fct_shipment_leg (leg_id, shipment_id, origin_location_id, destination_location_id, direction, tendered_at_utc, shipped_units), quantify how order variability grows upstream. fct_shipment_leg carries no sku_id and no link back to order lines, so shipped_units is the node's whole mixed-SKU flow and the measurement is node-level by construction. For each node compute weekly units flowing out (what it served downstream) and weekly units flowing in (what it ordered up), then the ratio of their variances on detrended, deseasonalised residuals. Walk parent_location_id from a store to echelon 0 and report the ratio at each hop, naming the hop where amplification is introduced. dim_location has no path column; build the chain yourself.

Approach
  1. Date orders by tendered_at_utc, not delivered_at_utc. Tender is the closest observable proxy for the moment the ordering decision was made; dating by delivery shifts the series by transit time and smears the variance you are trying to measure.
  2. Build both series on the same calendar weeks and in the same units. Variance scales with the aggregation window, so a ratio formed from daily inbound against weekly outbound measures the calendar, not the ordering rule. The units are mixed-SKU counts because the leg table carries no sku_id, so a node whose product mix drifts toward smaller or larger pack sizes moves both series for reasons unrelated to its ordering rule; print that caveat next to the number rather than implying a per-SKU result.
  3. Detrend and deseasonalise before taking variances, or compute the ratio on residuals from a simple weekly seasonal baseline. A growing node otherwise scores as amplifying, when all you have measured is its trend.
  4. Walk the parent chain iteratively: start from the store rows, join dim_location to itself on parent_location_id, and repeat until parent is null, capping the loop at the known echelon depth and asserting the path length matches the echelon difference so a data cycle raises rather than hangs.
  5. Read the output as a sequence, not a set of numbers. A pass-through node such as a cross-dock should sit near 1.0, and the hop where the ratio jumps is where a batching rule, a minimum order quantity or truckload rounding lives. That is the node to fix, not the node that is complaining.
Follow-up
  • The ratio at one hop is 4.2 and the node insists it orders exactly to forecast. What ordering rules would produce that number anyway?
  • What would it take to run this for a single SKU family rather than total units, given fct_shipment_leg carries neither a sku_id nor any link to order lines, and which hops would still be unmeasurable after you added that link?
  • What would you expect this measurement to look like during a promotion, and does that change your conclusion?

Cluster bootstrap for landed cost per delivered unit

mediumWorked solution
bootstrapclusteringratio estimatormix shift

You have delivered outbound legs from fct_shipment_leg: leg_id, shipment_id, leg_seq, carrier_id, origin_location_id, destination_location_id, shipped_units, cube_m3, linehaul_cost_cents, fuel_surcharge_cents, accessorial_cost_cents, expedite_premium_cents. In this extract the linehaul for a multi-stop load is booked entirely on leg_seq 1, so allocate it across the load's legs by cube first. Then estimate the difference in landed cost per delivered unit between two carriers on lanes both serve, with a 95 percent interval. Write the resampling yourself; no bootstrap library.

Approach
  1. Allocate the linehaul before anything else, in proportion to each leg's cube over the load's total cube, and keep the other three cost columns where they already sit. State the choice: allocating by weight instead reorders lanes whenever freight is bulky rather than dense.
  2. Restrict to the set of lanes both carriers actually serve in the window. Comparing over all lanes measures which carrier was assigned the cheap lanes, not which carrier is cheaper.
  3. Resample shipments, not legs. Legs of one load share a single allocated linehaul and a single dispatch decision, so treating them as independent draws understates the variance of the estimate.
  4. Recompute the statistic as a ratio of sums on every resample, total allocated cost over total shipped units. Taking the mean of leg-level cost-per-unit values instead gives a different estimand in which a one-unit leg counts as much as a full truckload.
  5. Report both the raw difference and a lane-standardised difference where lane weights are fixed at the pooled volume share, and say which one you would put in front of a decision maker.
Worked solution 35 min
  1. Compute per-shipment total cube, allocate the leg_seq 1 linehaul across legs by cube share, and form leg_total_cost from the four cost columns.
  2. Build the lane key from origin_location_id and destination_location_id, and keep only lanes with delivered volume from both carriers.
  3. Compute the point estimate per carrier as sum(leg_total_cost) / sum(shipped_units), and take the difference.
  4. Draw 2,000 bootstrap replicates by sampling shipment_ids with replacement within each carrier, rebuilding both sums from the sampled legs and recomputing the ratio difference.
  5. Take the 2.5th and 97.5th percentiles of the replicate differences, then repeat the whole procedure with resampling stratified inside lane to produce the mix-standardised interval.
EXPECTED RESULTA point difference in cents per delivered unit with a 95 percent percentile interval from 2,000 replicates, reported twice: raw, and standardised to a fixed lane mix. The two differ whenever the carriers' lane mixes differ, and can differ in sign.
Follow-up
  • The interval crosses zero. What would you need in volume or in window length to resolve a difference of 2 cents per unit?
  • One carrier's accessorials are rising while its linehaul is flat. What is that signature telling you about execution versus rates?
  • How does your interval change if one shipment accounts for 15 percent of the units?

For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Design one test end to end on paper
  • Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
  • Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
  • State in advance what you will do if the primary metric is flat while a secondary metric is significant.

Deliverable: A one-page test design with a decision rule written before launch.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Power arithmetic until it is automatic
  • Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
  • Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
  • Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.

Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.

Practice prompt ↗Practice prompt ↗
03Variance and the unit-of-analysis problem
  • Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
  • Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
  • Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.

Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.

Practice prompt ↗Practice prompt ↗
04Validity threats you can actually test for
  • Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
  • Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
  • Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.

Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05When randomization is not available
  • Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
  • Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
  • List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.

Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.

Practice prompt ↗Practice prompt ↗
06The readout query
  • Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
  • Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
  • Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.

Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.

Practice prompt ↗Practice prompt ↗
07Present it to someone who will not read the appendix
  • Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
  • Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
  • Rewrite your opening line so the recommendation lands before any methodology.

Deliverable: A one-page readout whose first line is the recommendation.

Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.

Explain forecast uncertainty to a non-technical general manager

easy
communicating uncertaintyservice levelexecutive communication

A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.

Approach
  1. Convert the error band into the units the approver already thinks in: weeks of cover, pallets, or dollars at risk, not a percentage or a confidence interval.
  2. Answer the right-or-wrong question directly rather than deflecting: the forecast will be wrong, the size of the wrongness is what you measured, and the build is sized against that size.
  3. State precisely what the 97 percent buys: it is a cycle service level, the probability of not running out during one replenishment cycle, so roughly three cycles in a hundred see a stockout. It is not the share of units shipped from stock. Unit fill rate also depends on the replenishment quantity, and when that quantity is large relative to the standard deviation of lead-time demand, which is the normal case outside lot-for-lot ordering, the fill rate sits above the cycle service number, often above 99 percent at a 97 percent cycle target.
  4. Give the two-sided consequence in money: what the build costs to carry at the applicable cost of capital plus obsolescence risk on shelf-life items, against the margin at risk from the stockouts it prevents.
  5. Close with the decision you want and the trigger that would reverse it, for example a lag-7 bias check after four weeks that reopens the number.
Follow-up
  • The GM says just give me one number. What do you give, and what do you refuse to give?
  • How would your answer change if the items were frozen with a 90-day shelf life?

Defend a finding that the expedite program bought nothing

medium
defending findingscounterfactual reasoninglanded cost

Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.

Approach
  1. Restate his objection in its strongest form before answering it, because a counterfactual worsening is a legitimate argument and treating it as an excuse ends the conversation.
  2. Separate what you measured from what you claimed: the data show no detectable service gain, not that expedite has no effect, and the difference is the entire argument.
  3. Test his hypothesis with the data you already have rather than defending in the abstract: split legs by exception_code = 'customs_hold' and by lane, and compare expedited against non-expedited legs on the same lanes in the same weeks, since if expedite were holding the line the expedited lanes should show a service gap over comparable non-expedited ones.
  4. Concede what is true: without a holdout you cannot rule out a protective effect, and the pre-period is contaminated by the disruption, so the honest statement is an upper bound on the gain rather than a zero.
  5. Hold the part that survives: the spend is real, it is concentrated in a small set of lanes, and nobody set a decision rule for when a leg gets expedited, which is a controllable problem independent of the counterfactual.
  6. Name the next measurement: a lane-level staggered switch-off with a stated burn-in, and say what it would cost and how long it would take.
Follow-up
  • He offers to run the switch-off only on his two best lanes. Why is that a problem, and what do you counter with?
  • His director asks you for a yes or no on cutting the budget today. What do you say?

Scope a vague request about rising inventory

easy
scopingstakeholder managementinventory turns

A planning director opens with: inventory is up twelve percent on flat shipments, get me an analysis by Friday. You have fct_inventory_daily (inventory_date, sku_id, location_id, on_hand_qty, standard_cost_cents) and fct_order_line (shipped_qty, unit_cogs_cents, requested_ship_date). Nothing else is specified: not the window, not the comparison basis, not whether the twelve percent is units or value. State the three questions you would ask before writing any SQL, name the decision each answer changes, and describe the first cut you would run if nobody answers you before Friday.

Approach
  1. Ask what decision hangs on the answer: a buy stop, a write-off provision, a policy review and a board slide all need different cuts, and the director usually has one of them in mind.
  2. Pin the measurement before the cause: units or value, which two windows are being compared, and whether the denominator is average daily on-hand or a period-end snapshot, since a month-end denominator biases turns high by ten to twenty percent and can manufacture the whole movement.
  3. Ask which part of the network is in scope, because echelon matters: a build at echelon 2 ahead of a promotion and a build of blocked_qty at a plant are different problems with different owners.
  4. Commit to a default if no answer arrives: value at standard cost, averaged across every daily snapshot in both windows, cut by echelon, then by abc_class and xyz_class, then by lifecycle_status to separate phase_out stock from active cover.
  5. Say out loud what the first cut cannot settle, so the director is not surprised: it localises the build but does not attribute it to forecast bias, a lot-size change or a supplier pulling orders in.
Follow-up
  • The build is concentrated in one product family at two DCs. What are your next two queries, and what would make you stop calling it a planning problem?
  • The director wants the number by Friday and the cause by Friday. Which do you drop, and how do you say so?
  • 01

    A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.

  • 02

    Over two quarters, expedite_premium_cents on outbound legs in fct_shipment_leg rose from 0.4 to 1.3 percent of landed cost, while perfect order rate moved from 91.2 to 91.5 percent, inside the week-to-week spread. The transport lead who sponsored the expedite program disputes the finding in a review in front of his director, arguing that your window contains a port disruption that would have made service worse without the spend. Present the finding, say what you concede on the spot, say what you hold, and name the evidence that would change your conclusion.

  • 03

    A planning director opens with: inventory is up twelve percent on flat shipments, get me an analysis by Friday. You have fct_inventory_daily (inventory_date, sku_id, location_id, on_hand_qty, standard_cost_cents) and fct_order_line (shipped_qty, unit_cogs_cents, requested_ship_date). Nothing else is specified: not the window, not the comparison basis, not whether the twelve percent is units or value. State the three questions you would ask before writing any SQL, name the decision each answer changes, and describe the first cut you would run if nobody answers you before Friday.

PracHub interview preparation framework ↗
Is this an official WestRock interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at WestRock. Rounds and questions reflect what candidates have reported, not a process WestRock has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How long does the interview process typically take?

While it varies by team and location, the process is generally completed within a few weeks across three core stages. Stay in close contact with your recruiter for updates on your specific timeline.

PracHub interview research ↗
Is the technical interview focused on whiteboard coding or project discussion?

Expect a hybrid approach. You will likely be asked to discuss your past projects in detail, but you should also be prepared for technical questions regarding software packages and data manipulation methodologies.

PracHub interview research ↗
What is the best way to stand out?

The most successful candidates are those who can clearly link their technical skills to business outcomes. Focus on the impact your projects had on the company’s bottom line or operational efficiency.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.