This guide covers what a Data Scientist at global consulting firm is expected to do and how to prepare for the interview.
Technical Screens
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Deep-Dive Case Studies
reportedThis round runs as a working session, so part of what it decides is whether you are useful to think with. The interviewer will interrupt: a hint that the data you assumed does not exist, a challenge to your metric, a nudge toward a branch you skipped. Treating those as interference is the common failure. Reason out loud while your thinking is still provisional so there is something to react to, and when a redirect arrives, take it instead of defending the path you had already started down.
What to demonstrate
- Whether your reasoning is audible while it is still unsettled, or only after you have privately decided
- What you do with a hint: absorb it and adjust, or argue for the original route
- Whether your clarifying questions have answers that would change your approach, as opposed to filling silence
- Whether you can be wrong about something in the middle of the case and keep moving without restarting
How to prepare
- Run practice cases with a partner instructed to interrupt twice: once to remove a data source you assumed existed, once to reject the metric you chose. Practise absorbing both without going back to the start.
- Before each practice case, write down the clarifying questions you plan to ask, then check afterwards whether any answer actually changed what you did. Drop the ones that did not.
- Explain an analysis you already know well to someone outside the field and have them stop you at every point where the reasoning jumped a step.
Behavioral Assessments
reportedRounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.
What to demonstrate
- Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
- What you changed afterwards, stated as a check you now run rather than a lesson you now believe
- Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you
How to prepare
- Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
- Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
- Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
PracHub editorial advice for the preparation topics above.
Pooling margin, realisation or overrun across pricing models
Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.
Trending utilisation or revenue on work_date without accounting for timesheet backfill
Time entries are created days to weeks after the work happens, and the backfill tail often runs two to six weeks. A dashboard keyed on work_date therefore shows the most recent weeks as a decline that reverses on every refresh. The fix is either to hold the reporting window back past the observed backfill tail (measure the tail with the timesheet submission lag metric rather than guessing) or to report an as-of-entered_at snapshot so the series is internally consistent, and to state which one you used.
Ending an analysis without a recommendation or next step
Close with what you would do and what would change your mind, stated as a condition you can check later. If the evidence is genuinely inconclusive, recommend the specific next measurement and say what it costs in time or exposure.
Extrapolating a first-week lift inflated by novelty effects
Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Describe a research project you led where you encountered a significan…
Describe a research project you led where you encountered a significant roadblock; how did you overcome it?
Approach
- Say what the estimate is of, and over what population it generalises.
- Translate the result into the decision it informs, in one plain sentence.
- Write down the assumption the method needs before you use the method.
Follow-up
- Which assumption here is most likely to be violated in practice?
- What sample size would you need to detect an effect half this size?
How would you handle missing values in a large-scale CRM dataset befor…
How would you handle missing values in a large-scale CRM dataset before running a regression analysis?
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
Measure the timesheet backfill curve and pick a reporting cutoff
time_entries has work_date (date), entered_at (timezone-aware UTC timestamp), hours and status. Given a snapshot_date, restrict to work_date in [snapshot_date - 180 days, snapshot_date - 60 days] so every cohort is fully observed. For k = 0..45, compute F(k): the share of a work_date cohort's final hours that already existed as of work_date + k days, pooled across cohorts. Return the 46-point curve and the smallest k with F(k) >= 0.99. Some rows are entered before the work date; those lags are real, not errors.
Approach
- Compute lag = (entered_at converted to the reporting timezone and taken as a date) - work_date in whole days, then clip negative lags to 0 instead of dropping them; leave and planned time are routinely entered ahead of the work date and dropping them deflates the early curve.
- Take cohort totals as groupby(work_date).hours.sum() over the restricted window. These are final only because the window stops 60 days short of the snapshot, which is why the restriction is in the prompt.
- Build the numerator by summing hours per (work_date, lag), sorting by lag, taking a per-cohort cumsum, then reindexing each cohort onto the full 0..45 lag grid and forward-filling, so a cohort with no entries at a given lag holds its previous level rather than disappearing.
- Pool as sum(numerators) / sum(denominators) at each k, not as the mean of per-cohort shares. Holiday weeks are tiny cohorts and would otherwise carry the same weight as a full week.
- Read k* off the pooled curve and report F(45) with it: if F(45) is below about 0.995 the tail runs past the grid and k* is a lower bound, not the answer.
Worked solution 25 min
- Restrict rows to the [snapshot - 180d, snapshot - 60d] window and compute lag_days = (entered_at.dt.tz_convert(tz).dt.normalize().dt.date - work_date).dt.days, then lag_days = lag_days.clip(lower=0).
- cohort_total = df.groupby('work_date').hours.sum(); by_lag = df.groupby(['work_date','lag_days']).hours.sum().
- Reindex by_lag onto MultiIndex.from_product([cohorts, range(0,46)]), fill 0, cumsum within work_date to get hours_by_k.
- F = hours_by_k.groupby(level='lag_days').sum() / cohort_total[cohorts_in_grid].sum(); assert F is non-decreasing.
- k_star = int(F[F >= 0.99].index.min()) if any, else report 'not reached within 45 days' along with F(45).
Follow-up
- The dashboard refreshes daily. Would you hold the window back past k*, or publish an as-of-entered_at series instead, and what does each choice cost the reader?
- One practice area has a tail twice as long as the rest. Does that change the firm-wide cutoff, or does it change what you publish per practice area?
How do you optimize a query that is performing poorly on a multi-billi…
How do you optimize a query that is performing poorly on a multi-billion row table?
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Describe how you would join multiple tables to build a feature set for…
Describe how you would join multiple tables to build a feature set for a churn prediction model.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a query using window functions to calculate a moving average of …
Write a query using window functions to calculate a moving average of customer spend over the last 30 days.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Deduplicate to the latest issued invoice line per engagement
fct_invoice_line holds invoice_line_id, invoice_id, engagement_id, line_type, period_end_date, amount_usd, issued_at and status; draft lines have issued_at NULL. Return exactly one row per engagement_id, its most recently issued line, with invoice_line_id, line_type, amount_usd and issued_at. Two lines on the same engagement can share an issued_at value to the second, so the result must be identical on every run against unchanged data. Write it with a window function, then say in one sentence when DISTINCT ON would be the better choice.
Approach
- Filter to issued lines (issued_at IS NOT NULL, status <> 'draft') before ranking. Ranking first and filtering afterwards can leave an engagement whose top-ranked row was a draft with no row at all in the output.
- Rank with ROW_NUMBER() OVER (PARTITION BY engagement_id ORDER BY issued_at DESC, invoice_line_id DESC). The surrogate key as the final ORDER BY term is what makes ties resolve the same way on every execution; without it the planner is free to return either tied row.
- Wrap the ranked query in a CTE or subquery and filter rn = 1, since a window function cannot appear in the WHERE clause of the query that computes it.
- Name why RANK() is wrong here: on tied issued_at it assigns 1 to both rows, so the result carries two rows for one engagement and any downstream SUM double counts.
- Compare against DISTINCT ON (engagement_id) ... ORDER BY engagement_id, issued_at DESC, invoice_line_id DESC: same answer, one sort, usually cheaper, but the window form generalises the moment you also want rn <= 3 or a second ranking in the same pass.
Worked solution 15 min
- Count distinct engagement_id over the filtered (issued, non-draft) line set. This is your target row count.
- Write the ROW_NUMBER CTE with both ORDER BY terms and select rn = 1.
- Assert the contract with GROUP BY engagement_id HAVING COUNT(*) > 1 over the result; it must return nothing.
- Deliberately construct or find an engagement with two lines sharing issued_at and confirm exactly one survives.
- Rewrite with DISTINCT ON and diff the two result sets row by row.
Follow-up
- Now return the latest line per engagement per line_type. What changes in the PARTITION BY, and what does that do to the row count?
- Return the latest and second-latest line together. Which of the two constructs survives that requirement?
- issued_at is TIMESTAMPTZ and finance reports on a single local calendar. Where can 'latest' flip between the two conventions?
If we observed a sudden 10% drop in daily active users on a client’s p…
If we observed a sudden 10% drop in daily active users on a client’s platform, how would you investigate the root cause?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you design a metric to track the success of a new customer l…
How would you design a metric to track the success of a new customer loyalty program?
Approach
- Fix the population and the time window before naming any metric.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you determine if an increase in a key metric is statistically s…
How do you determine if an increase in a key metric is statistically significant versus just noise?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
What are the common pitfalls when designing an A/B test for a marketin…
What are the common pitfalls when designing an A/B test for a marketing campaign?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Define competitive win rate for a quarterly pipeline review
Pipeline reporting currently shows win rate as won proposals divided by all proposals created. Using fct_proposal (proposal_id, expected_value_usd, stage, loss_reason, is_competitive, created_at, submitted_at, decided_at), write a definition you would defend to a finance lead: numerator, denominator, window and every exclusion. Decide whether to weight by count or by expected_value_usd, and justify the choice. State what happens to proposals still in 'qualifying', 'scoping' or 'submitted' at the snapshot date, and why dropping them is not a neutral act.
Approach
- Cohort on decided_at, not created_at. A created_at cohort mixes decided and undecided proposals, so the rate for a recent quarter is computed over whichever deals resolved fastest and drifts upward on every refresh.
- Set the denominator to stage IN ('won','lost'). Excluding 'withdrawn' and 'no_decision' is defensible only if you report their count and value alongside, because they are not missing at random: they skew toward deals that were heading for a loss.
- Filter is_competitive = TRUE. Sole-sourced follow-on work closes at near certainty and mechanically inflates the rate, so it belongs on its own line rather than inside the headline number.
- Report count-weighted and value-weighted side by side. They diverge when large proposals lose more often than small ones, which is the diagnostically useful case; the value-weighted version is the one that ties to revenue.
- Treat open proposals as right-censored rather than absent. Either restrict to cohorts old enough that most proposals have decided, picking that horizon off a Kaplan-Meier cycle-time curve, or publish the rate with the undecided count and value attached to it.
Worked solution 25 min
- Compute the naive rate (won / all created) and the proposed rate (won / won+lost, is_competitive = TRUE, cohorted on decided_at) for the same four quarters.
- For each quarter, count and sum the value of proposals in 'withdrawn' and 'no_decision', and of proposals still open at the snapshot.
- Compute the value-weighted version and compare it with the count-weighted one.
- Re-run the most recent quarter as it would have looked 30 days ago and record how much each definition moved.
Follow-up
- Win rate rose four points while proposal volume fell by a third. Which do you report first, and what single query separates a genuine quality shift from a volume shift?
- How would you detect partners reclassifying likely losses as 'withdrawn' to protect the rate?
Win rate jumped after a pipeline hygiene push
Competitive win rate by value jumped from 38% to 51% in one quarter. The metric sums expected_value_usd from fct_proposal where stage = 'won' over the same sum where stage IN ('won','lost'), filtered to is_competitive = TRUE and decided_at in the window. That quarter, partners were asked to clear stale opportunities out of the pipeline. You also have fct_proposal_stage_event. Quantify how much of the jump is selling and how much is the denominator changing shape, and state what you would monitor from now on.
Approach
- Stop looking at the ratio and plot its parts: won value and lost value by decision week, plus counts and value landing in the excluded stages 'withdrawn' and 'no_decision'. A ratio move with a flat numerator is a denominator story.
- Date the intervention and look for a spike in excluded-stage terminations around it. Those deals are not missing at random: opportunities that go stale and get cleaned up are disproportionately ones that were heading for a loss.
- Reconstruct history from fct_proposal_stage_event rather than reading the mutable fct_proposal.stage column, since the cleanup overwrote the current stage in place. Identify proposals that moved from 'submitted' or 'scoping' to 'withdrawn' after the push, and how long they had been open.
- Bound the effect: recompute win rate with withdrawn and no_decision counted as losses. The truth lies between that lower bound and the reported figure; report both numbers rather than picking one.
- Check the second lever: is_competitive is also mutable, so reclassifying a deal as sole-sourced removes it from the rate entirely. Count flips of that flag in the window.
- Propose the standing monitor: share of decided value terminating in excluded stages, published beside win rate, so the denominator can never move silently again.
Follow-up
- Some withdrawals are genuine, for example a client cancelling a procurement. How would you separate those from cleanup?
- What would you have asked for before the hygiene push to make this measurable afterwards?
- Does the same exclusion logic distort proposal cycle time, and in which direction?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Data people depend on systems owned by other teams, and much of the job is negotiating for instrumentation, access, or a fix to a broken pipeline. Prepare an example of getting something changed upstream that you did not control. Describe what you asked for, what you traded, and how you worked while you waited.
How do you handle a situation where a client disagrees with your data-…
How do you handle a situation where a client disagrees with your data-driven recommendation?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Your utilisation dashboard caused a staffing decision on incomplete data
Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.
Approach
- Name the probe: whether you own a reporting-design error rather than reclassifying it as someone else's timesheet compliance problem, and whether you fix the class of bug instead of the single week.
- Quantify before explaining. Measure the backfill curve directly from entered_at: for work_date D, the share of final hours that existed as of D plus k, for k from 1 to 45. That gives an observed tail length instead of a guessed cutoff.
- Change the artifact so the incomplete region cannot be read as a trend. Either end the trended series at snapshot minus the measured tail, or publish an as-of-entered_at series that is internally consistent, and label which one is on screen.
- Tell the person who acted, first and directly, with the corrected series and the specific decision to revisit. A correction that arrives after they notice costs more than the original error.
- Add a standing completeness tile: timesheet submission lag, the share of entries where entered_at date minus work_date exceeds seven days, so the dashboard shows its own reliability rather than depending on you remembering.
Follow-up
- Leadership still wants to see the current week. What do you show, and how do you label it?
- One practice runs a three-day lag and another twenty days. Do you set one firm-wide cutoff or one per practice, and what does that cost in comparability?
Explaining a censored days-to-pay number on a board slide
A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.
Approach
- Name the probe: whether you can give a non-technical decision-maker one number, the direction of the error in the alternative, and the reason, without teaching survival analysis.
- Explain the mechanism in business language. Invoices that have been paid are disproportionately the ones that pay fast; the slow ones are still open and therefore missing from the 38-day average. The error is one-directional and it grows as collections get worse, which is exactly when the number matters.
- Commit to one headline. Either the Kaplan-Meier median, or a restricted mean days-to-pay capped at a fixed horizon such as 90 days, which is easier for a finance audience to audit. Say which you used and why, and do not put both headline numbers on the slide.
- Make the chart the survival curve or a simple share-paid-by-day-k curve rather than a bar of averages, because the audience question is really when cash arrives, not a single moment.
- State the forward behaviour before you are asked: the figure for a recent quarter will rise or fall as open invoices resolve, so the slide carries the snapshot date and the share of invoices still open.
Follow-up
- The exec asks why the number printed on last quarter's slide no longer reproduces. What is your answer, and what would have prevented the question?
- How would you report this by account_tier without putting six survival curves on one slide?
- 01
How do you handle a situation where a client disagrees with your data-driven recommendation?
- 02
Six weeks ago you shipped a weekly billable-utilisation dashboard keyed on fct_time_entry.work_date. It showed a six-point drop across the three most recent weeks. A practice lead pulled two consultants off an engagement in response. The drop reversed on the next refresh, because time entries are created days to weeks after the work happens and entered_at trails work_date. Describe what you do now: the diagnosis, the change to the artifact, and the conversation with the person who acted on your number.
- 03
A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.
Is this an official global consulting firm interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at global consulting firm. Rounds and questions reflect what candidates have reported, not a process global consulting firm has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
It varies by location and team, but usually spans several weeks from the initial screen to the final round. Expect a high density of interviews in the final stages.
PracHub interview research ↗Is there a specific emphasis on coding?
Yes, but the focus is on practical data manipulation and SQL. You should be comfortable writing clean, efficient code to solve real-world data problems.
PracHub interview research ↗How much of the interview is behavioral?
You should expect at least one round dedicated to your background, leadership experiences, and how you handle conflict. Treat this with the same seriousness as the technical rounds.
PracHub interview research ↗What differentiates a top-tier candidate?
The most successful candidates are those who can balance technical depth with the ability to explain the business impact of their work. Being able to "think like a consultant" is a major differentiator.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22