global consulting firm · Data Scientist
Updated · 2026-09-24

global consulting firm Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

This guide covers what a Data Scientist at global consulting firm is expected to do and how to prepare for the interview.

Nearly every loop contains a round whose deliverable is a recommendation to someone non-technical. Practise stating a conclusion, the confidence attached to it, and the cost of being wrong in each direction, because that triple is the artifact being graded.

global consulting firm candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Compute utilisation against a defended availability denominatorCorrect late-entered timesheets before trending recent weeksReconstruct pipeline stage history from mutable rows

29 min read

Practice 16 Data Scientist prompts
16Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

This guide covers what a Data Scientist at global consulting firm is expected to do and how to prepare for the interview.

01

Technical Screens

reported

Much of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.

What to demonstrate

  • Whether the query you write matches the plan you just described
  • What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
  • Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly

How to prepare

  • Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
  • Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
  • Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
PracHub interview research ↗
02

Deep-Dive Case Studies

reported

This round runs as a working session, so part of what it decides is whether you are useful to think with. The interviewer will interrupt: a hint that the data you assumed does not exist, a challenge to your metric, a nudge toward a branch you skipped. Treating those as interference is the common failure. Reason out loud while your thinking is still provisional so there is something to react to, and when a redirect arrives, take it instead of defending the path you had already started down.

What to demonstrate

  • Whether your reasoning is audible while it is still unsettled, or only after you have privately decided
  • What you do with a hint: absorb it and adjust, or argue for the original route
  • Whether your clarifying questions have answers that would change your approach, as opposed to filling silence
  • Whether you can be wrong about something in the middle of the case and keep moving without restarting

How to prepare

  • Run practice cases with a partner instructed to interrupt twice: once to remove a data source you assumed existed, once to reject the metric you chose. Practise absorbing both without going back to the start.
  • Before each practice case, write down the clarifying questions you plan to ask, then check afterwards whether any answer actually changed what you did. Drop the ones that did not.
  • Explain an analysis you already know well to someone outside the field and have them stop you at every point where the reasoning jumped a step.
PracHub interview research ↗
03

Behavioral Assessments

reported

Rounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.

What to demonstrate

  • Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
  • What you changed afterwards, stated as a check you now run rather than a lesson you now believe
  • Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you

How to prepare

  • Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
  • Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
  • Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Pooling margin, realisation or overrun across pricing models

Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.

02

Trending utilisation or revenue on work_date without accounting for timesheet backfill

Time entries are created days to weeks after the work happens, and the backfill tail often runs two to six weeks. A dashboard keyed on work_date therefore shows the most recent weeks as a decline that reverses on every refresh. The fix is either to hold the reporting window back past the observed backfill tail (measure the tail with the timesheet submission lag metric rather than guessing) or to report an as-of-entered_at snapshot so the series is internally consistent, and to state which one you used.

03

Ending an analysis without a recommendation or next step

Close with what you would do and what would change your mind, stated as a condition you can check later. If the evidence is genuinely inconclusive, recommend the specific next measurement and say what it costs in time or exposure.

04

Extrapolating a first-week lift inflated by novelty effects

Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

13 technical prompts3 include a worked solution

Describe a research project you led where you encountered a significan…

medium
statistics and probability

Describe a research project you led where you encountered a significant roadblock; how did you overcome it?

Approach
  1. Say what the estimate is of, and over what population it generalises.
  2. Translate the result into the decision it informs, in one plain sentence.
  3. Write down the assumption the method needs before you use the method.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • What sample size would you need to detect an effect half this size?

How would you handle missing values in a large-scale CRM dataset befor…

medium
machine learning and modelling

How would you handle missing values in a large-scale CRM dataset before running a regression analysis?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Pick an evaluation metric that matches the cost of each error type, not a default.
  3. Check what information would not exist at prediction time, and exclude it.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • What would you monitor after launch to know the model is still valid?

Measure the timesheet backfill curve and pick a reporting cutoff

easyWorked solution
late-arriving datadata qualitycohort curves

time_entries has work_date (date), entered_at (timezone-aware UTC timestamp), hours and status. Given a snapshot_date, restrict to work_date in [snapshot_date - 180 days, snapshot_date - 60 days] so every cohort is fully observed. For k = 0..45, compute F(k): the share of a work_date cohort's final hours that already existed as of work_date + k days, pooled across cohorts. Return the 46-point curve and the smallest k with F(k) >= 0.99. Some rows are entered before the work date; those lags are real, not errors.

Approach
  1. Compute lag = (entered_at converted to the reporting timezone and taken as a date) - work_date in whole days, then clip negative lags to 0 instead of dropping them; leave and planned time are routinely entered ahead of the work date and dropping them deflates the early curve.
  2. Take cohort totals as groupby(work_date).hours.sum() over the restricted window. These are final only because the window stops 60 days short of the snapshot, which is why the restriction is in the prompt.
  3. Build the numerator by summing hours per (work_date, lag), sorting by lag, taking a per-cohort cumsum, then reindexing each cohort onto the full 0..45 lag grid and forward-filling, so a cohort with no entries at a given lag holds its previous level rather than disappearing.
  4. Pool as sum(numerators) / sum(denominators) at each k, not as the mean of per-cohort shares. Holiday weeks are tiny cohorts and would otherwise carry the same weight as a full week.
  5. Read k* off the pooled curve and report F(45) with it: if F(45) is below about 0.995 the tail runs past the grid and k* is a lower bound, not the answer.
Worked solution 25 min
  1. Restrict rows to the [snapshot - 180d, snapshot - 60d] window and compute lag_days = (entered_at.dt.tz_convert(tz).dt.normalize().dt.date - work_date).dt.days, then lag_days = lag_days.clip(lower=0).
  2. cohort_total = df.groupby('work_date').hours.sum(); by_lag = df.groupby(['work_date','lag_days']).hours.sum().
  3. Reindex by_lag onto MultiIndex.from_product([cohorts, range(0,46)]), fill 0, cumsum within work_date to get hours_by_k.
  4. F = hours_by_k.groupby(level='lag_days').sum() / cohort_total[cohorts_in_grid].sum(); assert F is non-decreasing.
  5. k_star = int(F[F >= 0.99].index.min()) if any, else report 'not reached within 45 days' along with F(45).
EXPECTED RESULTA monotone non-decreasing 46-point series F(0)..F(45) bounded by 1.0, plus an integer k* (or an explicit 'not reached by day 45') reported together with the value of F(45).
Follow-up
  • The dashboard refreshes daily. Would you hold the window back past k*, or publish an as-of-entered_at series instead, and what does each choice cost the reader?
  • One practice area has a tail twice as long as the rest. Does that change the firm-wide cutoff, or does it change what you publish per practice area?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Data people depend on systems owned by other teams, and much of the job is negotiating for instrumentation, access, or a fix to a broken pipeline. Prepare an example of getting something changed upstream that you did not control. Describe what you asked for, what you traded, and how you worked while you waited.

How do you handle a situation where a client disagrees with your data-…

medium
behavioural and stakeholder questions

How do you handle a situation where a client disagrees with your data-driven recommendation?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. State the situation in two sentences and spend the rest on your reasoning.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What would you do differently if you ran that project again?
  • How did you know the outcome was caused by your change?

Your utilisation dashboard caused a staffing decision on incomplete data

medium
late-arriving datapostmortemreporting windows

Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.

Approach
  1. Name the probe: whether you own a reporting-design error rather than reclassifying it as someone else's timesheet compliance problem, and whether you fix the class of bug instead of the single week.
  2. Quantify before explaining. Measure the backfill curve directly from entered_at: for work_date D, the share of final hours that existed as of D plus k, for k from 1 to 45. That gives an observed tail length instead of a guessed cutoff.
  3. Change the artifact so the incomplete region cannot be read as a trend. Either end the trended series at snapshot minus the measured tail, or publish an as-of-entered_at series that is internally consistent, and label which one is on screen.
  4. Tell the person who acted, first and directly, with the corrected series and the specific decision to revisit. A correction that arrives after they notice costs more than the original error.
  5. Add a standing completeness tile: timesheet submission lag, the share of entries where entered_at date minus work_date exceeds seven days, so the dashboard shows its own reliability rather than depending on you remembering.
Follow-up
  • Leadership still wants to see the current week. What do you show, and how do you label it?
  • One practice runs a three-day lag and another twenty days. Do you set one firm-wide cutoff or one per practice, and what does that cost in comparability?

Explaining a censored days-to-pay number on a board slide

easy
censoringkaplan-meierexecutive communication

A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.

Approach
  1. Name the probe: whether you can give a non-technical decision-maker one number, the direction of the error in the alternative, and the reason, without teaching survival analysis.
  2. Explain the mechanism in business language. Invoices that have been paid are disproportionately the ones that pay fast; the slow ones are still open and therefore missing from the 38-day average. The error is one-directional and it grows as collections get worse, which is exactly when the number matters.
  3. Commit to one headline. Either the Kaplan-Meier median, or a restricted mean days-to-pay capped at a fixed horizon such as 90 days, which is easier for a finance audience to audit. Say which you used and why, and do not put both headline numbers on the slide.
  4. Make the chart the survival curve or a simple share-paid-by-day-k curve rather than a bar of averages, because the audience question is really when cash arrives, not a single moment.
  5. State the forward behaviour before you are asked: the figure for a recent quarter will rise or fall as open invoices resolve, so the slide carries the snapshot date and the share of invoices still open.
Follow-up
  • The exec asks why the number printed on last quarter's slide no longer reproduces. What is your answer, and what would have prevented the question?
  • How would you report this by account_tier without putting six survival curves on one slide?
  • 01

    How do you handle a situation where a client disagrees with your data-driven recommendation?

  • 02

    Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.

  • 03

    A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.

PracHub interview preparation framework ↗
Is this an official global consulting firm interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at global consulting firm. Rounds and questions reflect what candidates have reported, not a process global consulting firm has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How long does the interview process typically take?

It varies by location and team, but usually spans several weeks from the initial screen to the final round. Expect a high density of interviews in the final stages.

PracHub interview research ↗
Is there a specific emphasis on coding?

Yes, but the focus is on practical data manipulation and SQL. You should be comfortable writing clean, efficient code to solve real-world data problems.

PracHub interview research ↗
How much of the interview is behavioral?

You should expect at least one round dedicated to your background, leadership experiences, and how you handle conflict. Treat this with the same seriousness as the technical rounds.

PracHub interview research ↗
What differentiates a top-tier candidate?

The most successful candidates are those who can balance technical depth with the ability to explain the business impact of their work. Being able to "think like a consultant" is a major differentiator.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.