As a Data Scientist at Accenture, you sit at the crucial intersection of advanced analytics, enterprise consulting, and scalable technology delivery. You will work within diverse client environments, ranging from consumer products and life sciences to global technology operations, turning complex data assets into actionable business strategies. Your daily work directly influences high-stakes decisions, helping global organizations modernize their data infrastructure, adopt emerging AI solutions, and optimize business operations through rigorous quantitative modeling.
This role requires a rare blend of deep technical execution and sophisticated stakeholder communication. You will not only build predictive models and analyze large-scale consumer datasets using modern tools like Databricks and Python, but you will also translate those technical findings into compelling narratives for executive leadership. Whether you are forecasting enterprise trends or designing advanced analytics frameworks, your impact is measured by your ability to drive tangible value in complex, ambiguous business landscapes.
Preparing for this position means mastering both rigorous foundational data science and the consultative problem-solving expected by top-tier professional services firms. You will face multidisciplinary interviewers who test your ability to scope unstructured problems, write optimal code, and design reliable experiments. Meeting these standards positions you as a trusted advisor capable of delivering end-to-end data science solutions for enterprise clients.
Initial Screening
reportedWhoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.
What to demonstrate
- Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
- Whether your reason for wanting the role points at the work itself rather than the company's reputation
- Whether your language signals the level being screened for: what you decided yourself versus what you were handed
How to prepare
- Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
- Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
- Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
Technical Interview
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Case-Based Discussion
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Behavioral Interview
reportedRounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.
What to demonstrate
- Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
- What you changed afterwards, stated as a check you now run rather than a lesson you now believe
- Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you
How to prepare
- Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
- Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
- Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
Final Round
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
14 candidate reports. Individual accounts describe a particular role and hiring cycle.
Accenture Software Engineer interview: three virtual and office rounds
My journey included three interviews, with a format change near the office stage. The first two were virtual, and the third took place in an Accenture office after scheduling adjustments. Before the technical conversations, I had a recruiter screen that lasted about 30 to 45 minutes. Then I spoke with someone from a developer background for another 30 to 45 minutes. The final step was behavioral…
Read full experienceAccenture Software Engineer interview: JavaScript, React, and architecture
I started with a recruiter screen covering my experience, skills, current role, notice period, salary expectations, the position, and the project it would support. Then came two technical conversations. The first focused on JavaScript and React fundamentals, problem solving, scenario questions, and the projects I had worked on. The second went deeper into architecture and design: scalability, per…
Read full experienceAccenture Software Engineer interview: Python, Django, FastAPI, Pandas, and coding exercises
I had a role-specific technical interview that opened with my project experience. The interviewer then went into Python and its ecosystem: Python itself, Django, FastAPI, and Pandas. They also checked my understanding of OOP and API concepts. The technical discussion quickly became practical. I had multiple Python coding exercises focused on problem solving, where the point was to see how I reaso…
Read full experienceAccenture Software Engineer interview: resume discussion followed by a take-home assessment
I began with a structured early conversation. I introduced myself and went through my resume in detail, especially volunteer work and an internship. We discussed the programming languages and tools I had used, why I was interested in Accenture, and the company's different job levels and what comes with them. That made the call feel more like orientation than a rapid-fire screen. Afterward, I rece…
Read full experienceAccenture Consultant interview: cases, routing, and SAP FICO
One path moved quickly through three interviews: two case-focused rounds and a less formal fit conversation with someone from a prospecting team. It was easy to follow and did not feel overly formal. In a different experience, I was told the first interviewer had not reviewed my CV and thought I might fit another team better. They still gave me a case, but the way my application was being routed…
Read full experiencePracHub editorial advice for the preparation topics above.
Pooling margin, realisation or overrun across pricing models
Fixed-fee margin falls with hours worked; uncapped time-and-materials margin rises with hours worked; retainer margin depends on neither. A quarter in which the firm sells more fixed-fee work will show a margin change caused entirely by mix, not by delivery performance, and the aggregate can move in the opposite direction to every individual pricing model. Always stratify by fct_engagement.pricing_model before comparing periods, and report the mix shift alongside the within-stratum change.
Modelling win rate on proposals with a recorded outcome, using fields written after the decision
Two failures compound here. First, stage IN ('withdrawn','no_decision') is not missing at random: those are disproportionately deals that were going to be lost, so training on won-plus-lost only inflates apparent win rate and distorts the coefficients. Second, fields like engagement_id, final scope and revised pricing are populated after the outcome is known, so including them leaks the label and produces a model with excellent backtest accuracy and no forward value. Restrict features to values knowable at submitted_at.
Extrapolating a first-week lift inflated by novelty effects
Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.
Reaching for a model before the target metric exists
Before naming an algorithm, write down the label, the prediction time, and the action that changes when the score crosses a threshold. If you cannot say what decision the output drives, any modelling choice is guesswork dressed up as method.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you account for seasonality and external market shocks when ana…
How do you account for seasonality and external market shocks when analyzing historical time-series data?
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Sanity-check the answer against a simple bound or a simulated case.
- Write down the assumption the method needs before you use the method.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
When would you choose a simpler logistic regression over a complex gra…
When would you choose a simpler logistic regression over a complex gradient boosting machine for an enterprise deployment?
Approach
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- What would you monitor after launch to know the model is still valid?
- Where could label leakage enter this setup?
Cluster bootstrap a margin change across few, unequal accounts
engagements has engagement_id, client_id, pricing_model, fees_usd, cost_usd and period in {pre, post} around a rate-card change. There are about 35 client_ids across roughly 400 engagements, and the top five clients carry a large share of fees. Estimate the pre-to-post change in value-weighted gross margin within pricing_model, and give a 95% interval by resampling whole client_ids with replacement, B = 2000. Write the bootstrap yourself; no resampling library. Report the account-weighted estimate alongside and explain any divergence.
Approach
- Write the statistic as a pure function of a dataframe first. Within each pricing_model, margin = (sum(fees) - sum(cost)) / sum(fees) per period; the headline is the fee-weighted average of the within-model differences using fixed pre-period weights, so a shift in the mix of work sold cannot masquerade as a change in delivery.
- Resample clusters, not rows: draw 35 client_ids with replacement and concatenate all engagements of each drawn client. A client drawn twice contributes its rows twice under distinct pseudo-ids, which is what preserves within-account correlation instead of averaging it away.
- Recompute the statistic on each replicate and take a percentile interval from the 2.5th and 97.5th quantiles. With concentrated fees the replicate distribution is skewed, so a symmetric point plus or minus 1.96 times a standard error is wrong in the tail that matters.
- Compute the account-weighted version (mean across clients of each client's margin change) next to the value-weighted one, and attribute the divergence to concentration rather than to noise; if they disagree in sign, that fact is the finding.
- Name the regime honestly: at roughly 35 clusters both cluster-robust standard errors and the pairs cluster bootstrap under-cover, so state the fix you would run next, which is a wild cluster bootstrap with Rademacher weights or CR2 with t(G-1) critical values.
Worked solution 40 min
- stat(df): pivot fees and cost by (pricing_model, period), compute per-model margins, difference post minus pre, and weight the differences by each model's pre-period fee share; return one scalar.
- Pre-split the frame into a dict of client_id to its rows, so each replicate is a concat of 35 preselected blocks rather than a boolean filter over 400 rows.
- rng = np.random.default_rng(11); for b in range(2000): draw client ids with replacement, concat their blocks, call stat, store; assert the replicate's nunique cluster draw count is 35 every time.
- ci = np.quantile(reps, [0.025, 0.975]); compute the account-weighted estimate as the unweighted mean of per-client margin changes over clients present in both periods.
- Re-run a row-level bootstrap on the same data for contrast and record both interval widths, then leave-one-client-out to see whether one account drives the point estimate.
Follow-up
- Implement the wild cluster bootstrap and show at what cluster count its p-value separates from the naive one.
- One client is 30% of fees. Show the interval with and without that account and say which one you would present, and to whom.
- What pre-period check would make you willing to call this a causal effect of the rate card rather than a correlation?
How would you identify and handle missing or corrupted data points acr…
How would you identify and handle missing or corrupted data points across a large distributed database using Python?
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Optimize a slow-running SQL query that joins multiple enterprise custo…
Optimize a slow-running SQL query that joins multiple enterprise customer tables with millions of records.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Engagement margin without fan-out across two fact tables
fct_engagement holds engagement_id, pricing_model and contract_value_usd. fct_invoice_line holds engagement_id, line_type, amount_usd and status. fct_time_entry holds engagement_id, hours, cost_rate_usd, is_billable and status. For engagements closed last year, return fees (line_type IN ('fees','milestone','credit_note') and status <> 'draft'), delivery cost (SUM(hours * cost_rate_usd) over approved entries, billable and non-billable alike) and gross margin, grouped by pricing_model. A single SELECT joining the engagement to both fact tables and then aggregating produces wrong numbers: say what the error is, then write the correct query.
Approach
- Do the arithmetic out loud. An engagement with 40 qualifying invoice lines and 900 approved time entries yields 36,000 rows on the double join: SUM(amount_usd) comes back 900 times too large and SUM(hours * cost_rate_usd) 40 times too large. The two inflation factors are different, so the ratio moves too — per engagement the double join replaces the cost-to-fee ratio with (cost / fees) * (line count / entry count), so a true 30 percent margin reads as 1 - 0.7 * (40 / 900), about 97 percent.
- Know which shape of this bug actually survives review. The ratio is preserved only where an engagement carries the same number of qualifying invoice lines as approved time entries, and then only as a row-count-weighted margin, 1 - SUM(cost_i * k_i) / SUM(fees_i * k_i), which still differs from the true pooled margin unless k_i is constant across engagements. So you get either a total absurd enough to tempt someone into a scaling patch, or, on a population where the two counts happen to track each other, a believable-looking ratio sitting on fees and cost that are each wrong by orders of magnitude.
- Aggregate each fact table to engagement grain in its own CTE, then LEFT JOIN both onto fct_engagement. One row per engagement then holds by construction rather than by inspection.
- Include non-billable delivery hours in cost. Rework, unbilled travel and pursuit time on a live account are real fully-loaded cost, and excluding them flatters fixed-fee work specifically, which is the mix you most need to see clearly.
- Treat zero and NULL as different: an engagement with no invoice lines has undefined margin, not zero. Divide by NULLIF(fees, 0) and keep those engagements in a labelled bucket instead of letting an inner join hide them.
- Group by pricing_model before reporting any firm-wide figure, and publish the fee mix next to it. Fixed-fee margin falls as hours rise, uncapped time-and-materials margin does not, so a shift in what was sold moves the pooled number with no change in delivery at all.
Worked solution 30 min
- Compute standalone control totals: total fees over the filtered invoice-line set and total cost over the approved time-entry set, both for the closed-engagement population.
- Build fees_by_engagement and cost_by_engagement as separate CTEs at engagement grain.
- LEFT JOIN both onto fct_engagement, compute margin with NULLIF on the denominator, and group by pricing_model.
- Run the naive double join on a single engagement and confirm its fee total equals the true total times that engagement's time-entry count, and its cost total the true cost times its invoice-line count.
- Report the fee mix by pricing_model alongside the margin column.
Follow-up
- Decompose a period-over-period margin move into a within-pricing-model component and a between-model mix component, and show the two sum to the total.
- Where does not_to_exceed_usd change the time-and-materials picture, and how would you find engagements that crossed it?
- Roll this up to the account through parent_client_id. What must the recursive CTE guard against?
How would you design a new engagement metric for a retail analytics mo…
How would you design a new engagement metric for a retail analytics mobile application?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you approach investigating an unexpected 15% drop in daily a…
How would you approach investigating an unexpected 15% drop in daily active users over a weekend?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How would you measure the long-term value of introducing an AI-driven …
How would you measure the long-term value of introducing an AI-driven recommendation engine to a consumer products platform?
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Your primary conversion metric dropped following a new release; walk m…
Your primary conversion metric dropped following a new release; walk me through your step-by-step metric drop diagnosis.
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How do you determine the required sample size and minimum detectable e…
How do you determine the required sample size and minimum detectable effect before launching an experiment?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
How would you diagnose and correct for sample ratio mismatch in an onl…
How would you diagnose and correct for sample ratio mismatch in an online experiment?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- What would you do if you could not randomise at all?
Define success for a rate-card increase before any rollout
A 6% rate-card increase is proposed for one practice area next quarter. Before it ships, define what success means. Available: fct_proposal (expected_value_usd, stage, loss_reason, is_competitive, submitted_at, decided_at, practice_area), fct_engagement (pricing_model, contract_value_usd, practice_area, client_id) and fct_invoice_line (line_type, amount_usd, billed_hours). Name the primary metric, the guardrail that genuinely opposes it, the unit of analysis, and the numeric threshold that would reverse the decision. State how many account-quarters are needed to detect an effect large enough to change the decision, and what you do if that exceeds what the practice has.
Approach
- Make the primary realised revenue per billable hour, not the list rate. The increase can be conceded straight back through discount and show up as no change at all. Numerator: SUM(amount_usd) on line_type IN ('fees','milestone'). Denominator: SUM(billed_hours) over the same lines.
- Name the opposing guardrail: value-weighted competitive win rate in that practice area, with the share of losses carrying loss_reason = 'price' as the supporting diagnostic. Raising price lifts the primary and depresses the guardrail by construction, so the rollout pays off only if revenue per hour rises faster than volume falls.
- Derive both break-evens before looking at outcome data. Revenue is preserved while the volume loss stays under 6 / 106, roughly 5.7%. For margin, write p for realised price per billable hour and c for avoidable cost per billable hour, and require c < p so the baseline margin is positive. Setting post-change margin (1 - L) x (1.06p - c) equal to pre-change margin (p - c) gives 1 - L = (p - c) / (1.06p - c), so the break-even volume loss is L = 0.06p / (1.06p - c). That is strictly above 5.7% whenever c > 0, collapses to 5.7% at c = 0, and stays below 100% as long as c < p; a practice already below cost at the old rate is not a pricing question. It holds only where c is genuinely avoidable on the volume you lose; on a salaried bench that nobody stands down or resells, no cost is saved and the margin break-even falls back to the revenue one. Report both ends and name the assumption the threshold uses.
- Fix the unit and confront the identification problem. Assignment is at practice_area while the outcome sits at engagement or account, so cluster at client_id. With one treated practice area there is a single treated cluster and a conventional difference-in-differences is not identified; propose a synthetic control over the other practice areas with the donor-pool condition stated, or a pre-registered before-and-after whose parallel-trends check runs on the pre-period with a named failure action.
- Run the power calculation before proposing the rollout and be willing to conclude the practice cannot detect the effect. Competitive decisions in one practice area number in the tens per quarter, so a 3-point win-rate move needs many quarters. If the requirement exceeds what exists, the honest outputs are a staged rollout with a pre-committed stopping rule, or shipping without measurement and saying so, not a manufactured p-value.
Worked solution 40 min
- Compute realised revenue per billable hour by practice-quarter for eight pre-period quarters from fct_invoice_line fees and milestone lines over billed_hours.
- Compute value-weighted competitive win rate per practice-quarter, cohorted on decided_at, denominator stage IN ('won','lost'), is_competitive = TRUE.
- Derive the two break-even volume losses: 6 / 106 for revenue, and 0.06p / (1.06p - c) for margin using the practice's realised price and avoidable cost per billable hour, then recompute with c treated as unavoidable to get the other end of the range.
- Count competitive decisions per quarter in the practice and compute the minimum detectable win-rate change at 80% power from a two-proportion test, treating that as a floor because clustering raises the real requirement.
- Write the rule: threshold, horizon, stopping condition, and the action taken when the power calculation says the effect is undetectable.
Follow-up
- Realised revenue per hour rose 4% and win rate fell 2 points on 31 competitive decisions. What do you conclude, and what is your interval on the win-rate move?
- How would you separate a price-driven loss from a capability-driven one when loss_reason is entered by the partner who lost the deal?
Utilisation fell three straight weeks; test entry lag first
A weekly billable-utilisation dashboard keyed on fct_time_entry.work_date shows three consecutive declines, roughly 74% to 66%. Billable headcount, the leave calendar and the holiday calendar are unchanged. You have fct_time_entry (consultant_id, work_date, entered_at, approved_at, hours, is_billable, charge_code, status) and dim_consultant (consultant_id, fte_fraction, is_billable_role, hire_date, termination_date). Decide whether the decline is real or an artefact of when timesheets are entered. Deliver a corrected series, the reporting cutoff you would adopt, and the evidence for that cutoff.
Approach
- Freeze the denominator before touching the numerator: available hours = scheduled workdays x 8 x fte_fraction, minus approved leave hours and region public holidays, prorated for hire_date and termination_date inside the week. Confirm it is flat across the three weeks, so any movement must be numerator.
- Measure the backfill curve instead of guessing it: for each work_date D in a mature window, compute the share of that day's eventual hours whose entered_at <= D + k, for k = 1..45. Read off k*, the lag at which the curve reaches about 99%.
- Restate at equal maturity: recompute every week counting only entries with (entered_at::date - work_date) <= k*, and suppress any week younger than k* days. Weeks are then comparable because each is observed at the same age.
- Check the approval gate separately: the metric counts status = 'approved', so a backlog of 'submitted' or 'rejected' rows reproduces the same shape with no entry lag at all. Plot hours by status per week.
- If the equal-maturity series is flat, report an artefact and publish the cutoff; if it still falls, segment by charge_code (client_delivery versus bench, training, internal_project) and by practice_area to locate the real movement.
Follow-up
- The backfill tail itself lengthened this quarter. Is that a finding, and what would you do with it?
- Leadership needs a number for the current week today. What do you give them, and how do you label it?
- How would this analysis change if approval, not entry, were the slow step?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.
Give an example of a time you uncovered an unexpected insight in data …
Give an example of a time you uncovered an unexpected insight in data that altered the direction of a business strategy.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
- Close with what you would do differently, concretely.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Defending a fixed-fee overrun finding against the selling partner
Over the trailing four quarters, fixed_fee engagements in one practice area show a median scope overrun ratio of 1.34, computed as approved client_delivery hours in fct_time_entry divided by fct_engagement.contracted_hours, and inception-to-date gross margin nine points below the time_and_materials book. The partner who sold most of that work says the denominator is wrong because change orders were signed, and each change order is a separate row linked by prior_engagement_id. You have ten minutes with the practice lead. Present and defend the finding.
Approach
- Say out loud what is being probed: whether a specific methodological objection makes you recompute or makes you repeat yourself. Restate the metric as numerator, denominator and window before defending anything.
- Test the objection empirically instead of debating it. Rebuild the denominator by walking the prior_engagement_id chain with a recursive CTE, summing contracted_hours across the original SOW and every extension, and summing approved delivery hours over the same chain.
- Report both numbers and say what each one answers: the original-SOW denominator answers whether the team held the scope that was signed, the chained denominator answers whether the firm estimated the total work correctly. Neither is a trick; they are different questions.
- Keep the margin claim separate from the overrun claim, and keep it stratified. Fixed_fee margin falls with hours worked while uncapped time_and_materials margin does not, so the nine-point gap is only meaningful within pricing_model, and a shift in the fixed_fee share of fees is reported alongside it.
- Close with the decision and a falsifier: if the chained overrun is near 1.0, the problem is scope control on the original SOW, and the check is whether change orders repriced at the standard rate card or at the original blended rate.
Follow-up
- The chained denominator narrows the overrun to 1.11. Does your recommendation change, and which number goes in the practice review?
- How would you distinguish an estimation problem at sale time from a delivery problem during execution, using only these tables?
Three requests, one week, and the one you defer
Three requests arrive in one week. Finance needs days sales outstanding recomputed for a board meeting in four days. A partner wants a proposal win-rate model for a pursuit review in three weeks. Delivery wants a staffing forecast, with no date attached. You have one week of your own capacity and no analyst. Write the prioritisation you send back, the request you defer, and the message you send to the person whose request you defer.
Approach
- Name the probe: whether you prioritise on decision dates and reversibility, or on who asked most forcefully.
- Score each request on four things: what decision it changes, the date that decision is made, the cost of being late, and the cost of being wrong. A board figure has a hard date and a high cost of being wrong; a forecast with no date has neither.
- Reduce scope rather than dropping work. The win-rate model is the largest piece and the easiest to get wrong, because features written after the decision, such as engagement_id and revised pricing, leak the label, and because withdrawn and no_decision proposals are not missing at random. A two-day descriptive win-rate cut by is_competitive and loss_reason answers most of what a pursuit review needs, with the model scoped separately.
- Defer explicitly, with a date and a smaller substitute, rather than leaving a request to decay quietly. Silence is read as agreement and then as failure.
- Put the trade-off in writing so it can be overturned by someone with more context than you have, and say what you would drop if the deferred request becomes urgent.
Follow-up
- The partner escalates to the practice lead. What do you change, and what do you refuse to change?
- What evidence would make you drop the finance work instead?
- 01
Give an example of a time you uncovered an unexpected insight in data that altered the direction of a business strategy.
- 02
Over the trailing four quarters, fixed_fee engagements in one practice area show a median scope overrun ratio of 1.34, computed as approved client_delivery hours in fct_time_entry divided by fct_engagement.contracted_hours, and inception-to-date gross margin nine points below the time_and_materials book. The partner who sold most of that work says the denominator is wrong because change orders were signed, and each change order is a separate row linked by prior_engagement_id. You have ten minutes with the practice lead. Present and defend the finding.
- 03
Three requests arrive in one week. Finance needs days sales outstanding recomputed for a board meeting in four days. A partner wants a proposal win-rate model for a pursuit review in three weeks. Delivery wants a staffing forecast, with no date attached. You have one week of your own capacity and no analyst. Write the prioritisation you send back, the request you defer, and the message you send to the person whose request you defer.
Is this an official Accenture interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Accenture. Rounds and questions reflect what candidates have reported, not a process Accenture has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process at Accenture?
The interview loops are moderately rigorous, balancing technical evaluations with assessments of your consulting and communication skills. Success depends on your ability to combine strong coding and statistical fundamentals with clear, structured problem-solving.
PracHub interview research ↗How much preparation time should I plan for?
Most candidates benefit from 4 to 6 weeks of dedicated preparation. Use this time to refresh your SQL window functions, review experimental design principles, and practice articulating your past project experience using structured behavioral frameworks.
PracHub interview research ↗What differentiates successful candidates from others?
Successful candidates distinguish themselves by connecting technical solutions directly to business value. Rather than just writing correct code, they explain their trade-offs, address edge cases proactively, and communicate with the poise expected of trusted advisors.
PracHub interview research ↗What is the typical timeline from initial screen to offer?
The timeline varies depending on the specific department and location, but a typical process moves from an initial HR call to technical screens and final interviews over the course of several weeks. Maintaining open communication with your recruiter helps ensure smooth scheduling.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22