KPMG · Data Scientist
Updated · 2026-09-24

KPMG Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at KPMG, you operate at the intersection of advanced analytics, strategic consulting, and business transformation. You are not merely building models in a vacuum; you are translating complex data into actionable insights that solve high-stakes problems for KPMG’s diverse global clientele. Your work directly influences how organizations optimize operations, mitigate risk, and leverage emerging technologies to maintain a competitive edge.

How much statistics you need depends on the flavour of the seat. Experiment-facing work wants you deep enough to notice that repeated looks at accumulating data inflate the false positive rate of a fixed-sample test; modelling-facing work wants estimation and honest uncertainty intervals.

KPMG candidates report 2 rounds · ≈ 2-4 weeks. The stages below are what candidates describe, not a published process.

Separate bookings, recognised revenue and collected cashCluster inference at the account, not the engagementSegment margin by pricing model before comparing

28 min read

Practice 14 Data Scientist prompts
14Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at KPMG, you operate at the intersection of advanced analytics, strategic consulting, and business transformation. You are not merely building models in a vacuum; you are translating complex data into actionable insights that solve high-stakes problems for KPMG’s diverse global clientele. Your work directly influences how organizations optimize operations, mitigate risk, and leverage emerging technologies to maintain a competitive edge.

The role demands a rigorous balance between technical depth and business acumen. You will be expected to navigate ambiguous problem spaces, communicate technical findings to non-technical stakeholders, and deliver solutions that are both theoretically sound and commercially viable. Whether you are working on predictive modeling, time-series forecasting, or modern AI implementations, your contribution is critical to the firm’s reputation for delivering data-driven excellence.

The environment at KPMG is fast-paced and results-oriented. Be prepared to demonstrate not just your ability to code, but your ability to translate that code into clear business value.

01

Initial Screening Call

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Assessments

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Pooling margin, realisation or overrun across pricing models

Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.

02

Treating accounts as independent observations

Revenue is concentrated: a small number of client_ids typically carries a large share of fees, and engagements within one account share a partner, a rate card and a delivery team. Ordinary standard errors computed over engagements therefore understate uncertainty badly. Cluster at client_id, and with fewer than roughly 40 clusters use a wild cluster bootstrap or a CR2 correction, because cluster-robust standard errors are downward-biased in that regime and will manufacture significance that a replication will not reproduce.

03

Reporting a mean for a heavy-tailed metric without saying what it hides

For spend, session length or items per order, a small fraction of units carries most of the total, so the mean has a wide standard error and one account can move it. Fix the handling before you see the result: cap or winsorise at a pre-declared percentile, and report the median or the share above a threshold next to the mean. Capping changes the estimand, so say which question the capped number answers, and check how much of any difference comes from the top 0.1 percent of units.

04

Generalising beyond the population the sample actually supports

State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

11 technical prompts3 include a worked solution

What are the primary challenges when working with time-series data com…

medium
statistics and probability

What are the primary challenges when working with time-series data compared to cross-sectional data?

Approach
  1. Sanity-check the answer against a simple bound or a simulated case.
  2. Say what the estimate is of, and over what population it generalises.
  3. Translate the result into the decision it informs, in one plain sentence.
Follow-up
  • What sample size would you need to detect an effect half this size?
  • Which assumption here is most likely to be violated in practice?

Explain the application of Bayesian techniques in a business context.

medium
statistics and probability

Explain the application of Bayesian techniques in a business context.

Approach
  1. Say what the estimate is of, and over what population it generalises.
  2. Translate the result into the decision it informs, in one plain sentence.
  3. Write down the assumption the method needs before you use the method.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • What sample size would you need to detect an effect half this size?

How do you approach feature selection when dealing with high-dimension…

medium
statistics and probability

How do you approach feature selection when dealing with high-dimensional datasets?

Approach
  1. Write down the assumption the method needs before you use the method.
  2. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  3. Sanity-check the answer against a simple bound or a simulated case.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • What sample size would you need to detect an effect half this size?

What metrics would you use to evaluate the performance of a predictive…

medium
machine learning and modelling

What metrics would you use to evaluate the performance of a predictive model, and why?

Approach
  1. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  2. Say how the offline result would be validated online before it is trusted.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

Can you explain the trade-offs between different loss functions in reg…

medium
machine learning and modelling

Can you explain the trade-offs between different loss functions in regression tasks?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Set a baseline first, so any model has something honest to beat.
  3. Check what information would not exist at prediction time, and exclude it.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • What would you monitor after launch to know the model is still valid?

Build a dense consultant-week hours spine with zero weeks

easyWorked solution
calendar spinegroupbymissing rows

You are given time_entries (time_entry_id, consultant_id, engagement_id, charge_code, work_date as datetime64, hours, is_billable, status) and a list of consultant_ids in scope plus a start and end date. Produce a DataFrame with exactly one row per consultant per Monday-anchored week in that range, with columns consultant_id, week_start, billable_hours, total_hours. Weeks in which a consultant logged nothing must appear as zeros rather than be absent. Count only rows with status == 'approved'. Do not loop over consultants and do not use resample.

Approach
  1. Anchor weeks arithmetically: week_start = work_date - pd.to_timedelta(work_date.dt.weekday, unit='D'). Avoid dt.isocalendar().week, whose numbers restart each January and do not sort across a year boundary.
  2. Construct the spine from the inputs, not from the data: pd.date_range(start, end, freq='W-MON') crossed with the in-scope consultant_ids via pd.MultiIndex.from_product, so its size is fixed before any aggregation runs.
  3. Aggregate approved entries once with a single groupby on [consultant_id, week_start], producing total hours and billable hours (hours.where(is_billable, 0).sum()) in the same pass.
  4. Left-merge the aggregate onto the spine, fillna(0.0) on both hour columns, and keep float dtypes so later ratios do not integer-divide.
  5. Assert the shape before returning: len(out) == n_consultants * n_weeks, and out.total_hours.sum() equals the sum of hours over the filtered input inside the window.
Worked solution 20 min
  1. Filter to status == 'approved' and to work_date within [start, end], then derive week_start by subtracting the weekday offset.
  2. Build weeks = pd.date_range(start_monday, end, freq='W-MON') and spine = pd.MultiIndex.from_product([consultant_ids, weeks], names=['consultant_id','week_start']).to_frame(index=False).
  3. agg = filtered.assign(billable_hours=filtered.hours.where(filtered.is_billable, 0.0)).groupby(['consultant_id','week_start'], as_index=False)[['hours','billable_hours']].sum().
  4. out = spine.merge(agg, how='left', on=['consultant_id','week_start']).fillna({'hours':0.0,'billable_hours':0.0}).rename(columns={'hours':'total_hours'}).
  5. Sort by consultant_id then week_start and run the two assertions on row count and hour conservation.
EXPECTED RESULTA frame of exactly len(consultant_ids) x len(weeks) rows, sorted by consultant then week, in which total_hours sums to the approved hours inside the window and zero-hour weeks are present as 0.0 rather than missing.
Follow-up
  • Utilisation gets computed on this spine. What breaks for a consultant hired mid-window or terminated mid-window, and where would you fix it?
  • How would you extend this to consultant x engagement x week without the spine exploding to a row count nobody can hold in memory?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.

How do you handle imbalanced datasets in classification models?

medium
behavioural and stakeholder questions

How do you handle imbalanced datasets in classification models?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Close with what you would do differently, concretely.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

Defending your own impact claim without randomisation or clean units

hard
self-selectionclustered inferenceimpact measurement

At your review you plan to claim that the realisation dashboard you built recovered 1.4 million dollars. The evidence is that engagement-month realisation rose four points over two quarters among engagements whose leads used it. Adoption was voluntary. There are about sixty client accounts and the top five carry most fees. A new rate card shipped in the same quarter. Write the claim you can defend, the estimate you would actually produce, and what you say when asked for a causal number you cannot get.

Approach
  1. Name the probe: whether you can separate the number you want from the number the data supports, under review pressure, without either inflating it or retreating to saying nothing can be known.
  2. State both identification problems concretely. Voluntary adoption means adopting leads are plausibly the ones who already manage realisation, so the comparison is confounded at the person level. The rate card changes bill_rate_usd, which sits in the realisation denominator, so part of the four-point move is arithmetic rather than behavioural.
  3. Neutralise what you can. Recompute realisation with bill rates snapshotted on work_date, or hold the denominator at the old rate card, so the rate-card change cannot move the metric by construction. Then rerun the comparison.
  4. Get the inference right for the unit count. Cluster at client_id, not engagement, because engagements in one account share a partner, a rate card and a team. With sixty accounts and five carrying most fees, the effective cluster count is far below sixty, so report a wild cluster bootstrap interval rather than plain cluster-robust standard errors, which are biased downward in that regime.
  5. Report both weightings and explain the divergence: an account-weighted estimate describes the typical account, a value-weighted one describes the revenue, and if they disagree a small number of accounts is carrying the result. Then give the decision-relevant sentence: the defensible range, whether its lower bound still clears the build cost, and what a proper staggered rollout would have bought.
Follow-up
  • The pre-period trends for adopters and non-adopters are not parallel. What do you report then?
  • You get to design the next rollout. What do you change so the same question is answerable, without randomising individual accounts?

Explaining a censored days-to-pay number on a board slide

easy
censoringkaplan-meierexecutive communication

A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.

Approach
  1. Name the probe: whether you can give a non-technical decision-maker one number, the direction of the error in the alternative, and the reason, without teaching survival analysis.
  2. Explain the mechanism in business language. Invoices that have been paid are disproportionately the ones that pay fast; the slow ones are still open and therefore missing from the 38-day average. The error is one-directional and it grows as collections get worse, which is exactly when the number matters.
  3. Commit to one headline. Either the Kaplan-Meier median, or a restricted mean days-to-pay capped at a fixed horizon such as 90 days, which is easier for a finance audience to audit. Say which you used and why, and do not put both headline numbers on the slide.
  4. Make the chart the survival curve or a simple share-paid-by-day-k curve rather than a bar of averages, because the audience question is really when cash arrives, not a single moment.
  5. State the forward behaviour before you are asked: the figure for a recent quarter will rise or fall as open invoices resolve, so the slide carries the snapshot date and the share of invoices still open.
Follow-up
  • The exec asks why the number printed on last quarter's slide no longer reproduces. What is your answer, and what would have prevented the question?
  • How would you report this by account_tier without putting six survival curves on one slide?
  • 01

    How do you handle imbalanced datasets in classification models?

  • 02

    At your review you plan to claim that the realisation dashboard you built recovered 1.4 million dollars. The evidence is that engagement-month realisation rose four points over two quarters among engagements whose leads used it. Adoption was voluntary. There are about sixty client accounts and the top five carry most fees. A new rate card shipped in the same quarter. Write the claim you can defend, the estimate you would actually produce, and what you say when asked for a causal number you cannot get.

  • 03

    A finance lead wants average days to get paid for the last four quarters, for a board slide. In fct_invoice_line, rows with status in ('issued','partially_paid','disputed') have paid_at NULL. The mean of (paid_at - issued_at) over paid invoices is 38 days. A Kaplan-Meier median, right-censoring the open invoices at snapshot_date minus issued_at, is 51 days. You get two sentences and one chart. Explain the number you put on the slide, the gap between the two figures, and what will make that number move next quarter.

PracHub interview preparation framework ↗
Is this an official KPMG interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at KPMG. Rounds and questions reflect what candidates have reported, not a process KPMG has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult are the technical interviews?

They are considered demanding. You should be prepared for deep-dive questions that go beyond high-level theory and into the implementation details of your past projects.

PracHub interview research ↗
What is the company culture like?

KPMG values honesty and direct communication. You will find that team members are often straightforward about project challenges and expectations, including the intensity of the work.

PracHub interview research ↗
How long does the entire process take?

While it varies, the average process from the first call to an offer often takes around one month.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.