Accenture Federal Services · Data Scientist
Updated · 2026-09-24

Accenture Federal Services Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Accenture Federal Services, you play a pivotal role in helping the United States federal government solve complex operational, strategic, and security challenges. You bridge the gap between raw, multi-source data and actionable intelligence, designing analytical workflows that directly influence defense, national security, public safety, and civilian agency missions. Your day-to-C day work involves translating ambiguous mission requirements into robust data pipelines, machine learning models, and intuitive dashboards that empower stakeholders to make rapid, data-driven decisions.

How much statistics you need depends on the flavour of the seat. Experiment-facing work wants you deep enough to notice that repeated looks at accumulating data inflate the false positive rate of a fixed-sample test; modelling-facing work wants estimation and honest uncertainty intervals.

Accenture Federal Services candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Model censored case durations without survivorship biasEstimate eligible populations outside administrative dataSeparate obligations, outlays and ceilings correctly

36 min read

Practice 18 Data Scientist prompts
18Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Accenture Federal Services, you play a pivotal role in helping the United States federal government solve complex operational, strategic, and security challenges. You bridge the gap between raw, multi-source data and actionable intelligence, designing analytical workflows that directly influence defense, national security, public safety, and civilian agency missions. Your day-to-C day work involves translating ambiguous mission requirements into robust data pipelines, machine learning models, and intuitive dashboards that empower stakeholders to make rapid, data-driven decisions.

This role requires balancing technical execution with high-touch stakeholder collaboration across diverse federal client environments. You will work alongside data engineers, product managers, and domain specialists to extract insights from massive, secure datasets while maintaining strict data governance. Whether you are deploying advanced natural language processing models, fine-tuning large language models, or architecting custom analytics on cloud platforms like AWS and Databricks, your contributions directly impact the safety, efficiency, and effectiveness of government programs.

Succeeding in this role demands a unique combination of core technical competence and adaptability. You must be comfortable navigating secure infrastructure, writing optimized code in Python and SQL, and communicating complex technical concepts to non-technical leaders. Expect an environment where intellectual curiosity, rigorous quantitative analysis, and a commitment to public service are deeply valued and rewarded.

01

Recruiter Screening

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Screening

reported

A handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.

What to demonstrate

  • Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
  • Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
  • Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.

How to prepare

  • Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
  • Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
  • If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
PracHub interview research ↗
03

Core Interview Rounds

reported

Rounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.

What to demonstrate

  • Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
  • The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
  • Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options

How to prepare

  • Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
  • Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
  • Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
PracHub interview research ↗
04

Behavioral Interviews

reported

Most of the weight in this round sits on the disagreement questions. Data work routinely produces an answer someone senior did not want, and the interviewer is trying to learn what you do in that hour. Both failure modes are common: folding as soon as a director pushes back, and treating the pushback as ignorance to be corrected with a better chart. A strong answer usually contains a specific thing the other person knew that you did not, and describes how you found out whether it changed the conclusion.

What to demonstrate

  • Whether you can state the other side's argument accurately before you explain why you disagreed
  • What you treated as evidence during the disagreement, such as a rerun under their assumption or a holdout check, rather than persuasion technique
  • Whether you distinguish being overruled from being wrong, and can give an example of each

How to prepare

  • Write out one disagreement where you turned out to be wrong, and say what in the data misled you. Candidates prepare the story where they were right, and the follow-up asks for the other one.
  • For your main disagreement story, be ready to say what result would have made you drop your position. If no such result exists, you were not arguing from the data.
  • Practise stating the opposing position out loud in one sentence the stakeholder would accept, then continue the story.
PracHub interview research ↗
05

Technical Case Studies

reported

Underneath the business framing, this round is usually asking whether you can turn a fuzzy goal into a quantity that could be computed from data such a business would plausibly hold. That means a metric with a stated numerator, denominator, eligibility rule and time window, plus an honest account of the conditions under which it would mislead you. Answers come apart when a candidate names a familiar metric and never defines it, because every follow-up then lands on an ambiguity that was left open and the candidate has to invent the definition under pressure.

What to demonstrate

  • Whether a named metric arrives with its denominator, eligibility rule and window attached rather than assumed
  • Whether the measure follows from the mechanism you proposed, or is a recognisable metric retrofitted to it afterwards
  • Whether you name a guardrail that would reveal the gain came from somewhere you did not want it to come from
  • Whether you can say what data the plan requires and what you would settle for if that logging were never implemented

How to prepare

  • Take five metrics you reach for by reflex and write each as one sentence containing numerator, denominator, eligibility rule and time window. The ones you cannot finish are the ones that will fail under follow-up.
  • For a product you use daily, write the measurement plan you would propose for a change to it: primary metric, one guardrail, the unit of analysis, and the table the numbers would come from.
  • Practise the substitution question. For three metrics you like, write what you would measure instead if the event you depend on were not being logged.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Computing coverage or take-up with an administrative denominator, for example approved applications divided by all applications

That ratio measures throughput among people who already found the system, not delivery to the people entitled to the service. It gets better when outreach is cut, because the marginal applicant is the one most likely to be denied or to abandon, and it gets worse when a new access channel brings in harder cases. The correct denominator is a modelled eligible population from survey microdata run through the eligibility rules, and the gap between it and the applicant count is usually the finding.

02

Dividing an administrative count by a survey population estimate and reporting the rate as if it were exact

At block-group and tract scale the published margin of error is often a large fraction of the estimate, so the variance of the resulting rate is dominated by the denominator rather than the numerator, and a leaderboard of areas ranked by point estimate is largely a leaderboard of the smallest and noisiest areas. Propagate the denominator error into the rate, or aggregate up until the relative standard error is acceptable, and state the threshold used. The same join also breaks silently across boundary vintages: a geo_code can refer to a different physical area after a redraw, so joining current-vintage geography onto historical events reassigns records and manufactures a trend break.

03

Generalising beyond the population the sample actually supports

State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.

04

Analysing at a different unit than the one randomised

Say out loud what was randomised (user, device, account, cluster) and make the analysis unit match, or account for the clustering with cluster-robust standard errors, the delta method, or aggregation up to the randomised unit. Randomising users and then running a test over sessions understates variance and inflates the false-positive rate.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

15 technical prompts3 include a worked solution

Explain the false discovery rate and how you adjust p-values when cond…

medium
statistics and probability

Explain the false discovery rate and how you adjust p-values when conducting thousands of simultaneous anomaly detection checks.

Approach
  1. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  2. Say what the estimate is of, and over what population it generalises.
  3. Translate the result into the decision it informs, in one plain sentence.
Follow-up
  • What sample size would you need to detect an effect half this size?
  • Which assumption here is most likely to be violated in practice?

How do you test for homoscedasticity and normality in residuals when b…

medium
machine learning and modelling

How do you test for homoscedasticity and normality in residuals when building a predictive regression model for resource allocation?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Say how the offline result would be validated online before it is trusted.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

Publish on-time decision share anchored on the due date

easyWorked solution
metric definitioncensoringpandas

You are given applications as a DataFrame with application_id, program_code, submitted_at, sla_due_at, decision_at (NaT while open) and status, plus a scalar extract_date. Implement on-time decision share monthly, keyed on the month of sla_due_at: the denominator is every application whose sla_due_at falls in that month whether or not it has been decided, and the numerator is those with decision_at <= sla_due_at. Flag months that are not yet complete. Then report, for one recent month, the value you would have published had you keyed on decision_at over decided rows only.

Approach
  1. Restrict to rows with submitted_at populated. Drafts and abandoned applications never acquire an sla_due_at, and sweeping them in gives a denominator that moves with intake abandonment rather than with adjudication.
  2. Build the denominator by flooring sla_due_at to month and counting every row in that month, including rows where decision_at is NaT. Keeping undecided rows in the denominator is the entire reason the definition anchors on the due date, so this is the line to write first.
  3. Numerator is decision_at.notna() & (decision_at <= sla_due_at). NaT comparisons already evaluate False, but write the notna() explicitly so the intent survives someone later swapping in a fillna.
  4. Mark a month incomplete when its due dates extend past extract_date, since a case whose due date is still in the future can yet be decided on time. Publishing those as provisional stops the current month reading as a collapse.
  5. Recompute keyed on decision_at over decided rows only and publish it beside the primary series rather than subtracting one from the other. The two denominators are different sets, every application due in the month against every application decided in the month, so the month-level gap between the two shares is not a population and its sign is not predictable. Report open_past_due as its own column instead, since that count is what the decision-keyed version cannot see.
Worked solution 20 min
  1. Filter to submitted_at.notna() and sla_due_at.notna(), counting any submitted row that is missing a due date rather than letting it drop silently, and derive due_month = sla_due_at.dt.to_period('M').
  2. Group by due_month: denominator is the row count, numerator is the sum of (decision_at.notna() & (decision_at <= sla_due_at)).
  3. Add is_complete by comparing the month end against extract_date, and carry a column for open_past_due = decision_at.isna() & (sla_due_at < extract_date).
  4. Separately group decided rows by decision_at month and compute the same ratio over that base for the comparison column.
  5. Return one frame with both series and the open past-due count, so the gap between them is readable without a second query.
EXPECTED RESULTA monthly frame carrying due_month, denominator, on_time, share, open_past_due, is_complete and the decision-keyed comparison. Neither share dominates the other, because the denominators are disjoint sets rather than nested ones. A month in which staff clear a block of already-late older cases posts a decision-keyed share far below the due-date-keyed share for that same calendar month, since those late decisions enter the decision-keyed denominator and can never enter its numerator; a quiet month whose due cases are largely still open posts the reverse. Summed over all history the two numerators count exactly the same applications, and only the bucketing and the denominators differ.
Follow-up
  • A unit spends one week clearing 400 cases that were already past their due date. The decision-keyed share for that month drops hard and the due-date-keyed series does not move in any month. Explain both to a director in two sentences and say which number you would publish.
  • Decision events in the case log can be backdated by staff. Which clock should this metric use, and how would you measure whether the choice changes the published number?
  • The service standard itself changed mid-month under a new rule version, so sla_due_at is derived from two different standards inside one bucket. How do you report that month?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Have two ready. In one, the data was on your side and you had to move someone who outranked you. In the other, the pushback was correct and you changed position. The second is the harder story and it lands better, because it shows you separate being right from being attached to an answer. Name the person's actual objection.

How do you handle network effects and interference between treatment a…

medium
behavioural and stakeholder questions

How do you handle network effects and interference between treatment and control groups when testing a collaborative software feature?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. State the situation in two sentences and spend the rest on your reasoning.
  3. Close with what you would do differently, concretely.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Disagree with a proposal to auto-deny on unreturned evidence

medium
metric gamingdisparate impactrollout design

A product manager proposes auto-denying any application sitting in status 'pending_evidence' for 14 days, writing denial_reason_code = 'evidence_not_returned', and projects a large fall in median days to decision. You have fact_application, fact_case_event and dim_geography. Write a one-page disagreement: the three numbers you would compute before the change ships, what each would have to show for the proposal to be sound, and the result that would make you drop your objection.

Approach
  1. Name the mechanism plainly: a procedural denial removes a case from the pending pool without a substantive decision, so the timeliness improvement is arithmetic. Any argument that does not address this is a disagreement about values rather than about the proposal.
  2. Number one, appeal overturn rate split by procedural versus substantive denial_reason_code, on a denial cohort with a 180-day follow-up. If procedural denials are overturned or remanded far more often, the rule manufactures rework that lands in cost per sustained outcome rather than in the timeliness number.
  3. Number two, the distribution of the lag from evidence_requested to evidence_received in fact_case_event, split by channel and preferred_language_code. If the mass beyond 14 days is concentrated in mail and in-person intake or in non-English speakers, the threshold has a disparate impact before any model is involved, and comparing like with like means holding evidence_item_code fixed.
  4. Number three, the reapplication rate: how often an evidence_not_returned denial is followed by a new fact_application row from the same constituent_id within 90 days. Those are the same case counted twice and the throughput gain is illusory.
  5. State the falsifier explicitly. Flat overturn rates across codes, a thin tail past 14 days that does not vary by channel or language, and rare reapplication would make the rule reasonable, and you will say so in writing.
  6. Propose the cheaper decision: a staged rollout by adjudicator_unit_id with pre-registered guardrails and a stopping rule, rather than a blanket switch that cannot be evaluated afterwards.
Follow-up
  • The product manager says the published service standard requires the case to close. Does that change your position, and what do you ask for instead?
  • Design the staged rollout: which units, which comparison, and what result stops it.

Explain why a ranked district table should not be published

easy
small-area estimationuncertaintyexecutive communication

You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.

Approach
  1. Convert each published margin to a standard error first: at 90 percent confidence the standard error is moe divided by 1.645. Compute the relative standard error of the denominator for every area in the table.
  2. Show where the uncertainty lives. Treating the administrative count as fixed and independent of the survey estimate, the rate's relative standard error equals the denominator's, so an area with a 24 percent relative standard error on population has a 24 percent relative standard error on its rate no matter how clean fact_service_request is.
  3. Demonstrate the instability rather than asserting it: resample each denominator from its published standard error, recompute the ranking a few thousand times, and report how often each area actually lands in the worst twelve. Areas that appear in only a third of draws are not findings.
  4. Replace the rank with something that survives the uncertainty: aggregate to a geo_level where the relative standard error clears a threshold you state out loud, or group areas into tiers whose intervals do not overlap, and name the threshold as a choice you made rather than a standard.
  5. Give the executive one sentence she can repeat without you in the room, for example that the data supports naming a group of high-demand areas but not ordering them.
Follow-up
  • Which relative standard error threshold do you use, and why is it a judgement rather than a rule?
  • Two of the areas you aggregated sit across a boundary redraw. What breaks in the join, and how would you notice?
  • 01

    How do you handle network effects and interference between treatment and control groups when testing a collaborative software feature?

  • 02

    A product manager proposes auto-denying any application sitting in status 'pending_evidence' for 14 days, writing denial_reason_code = 'evidence_not_returned', and projects a large fall in median days to decision. You have fact_application, fact_case_event and dim_geography. Write a one-page disagreement: the three numbers you would compute before the change ships, what each would have to show for the proposal to be sound, and the result that would make you drop your objection.

  • 03

    You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.

PracHub interview preparation framework ↗
Is this an official Accenture Federal Services interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Accenture Federal Services. Rounds and questions reflect what candidates have reported, not a process Accenture Federal Services has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
What level of interview difficulty should I expect for the Data Scientist role?

The interview loops range from straightforward conversational rounds to rigorous technical assessments involving live coding, system design, and statistical problem-solving. Preparation should cover both foundational coding in Python and SQL as well as high-level product and experimental design principles.

PracHub interview research ↗
How important is an active security clearance for getting hired?

An active security clearance (such as Secret, Top Secret, or TS/SCI with polygraph) is often a mandatory prerequisite depending on the specific client engagement and team. Be sure to clarify clearance requirements with your recruiter early in the screening process.

PracHub interview research ↗
How long does the typical interview process take from initial screen to offer?

The timeline varies based on scheduling availability and clearance verification, but typically spans several weeks from the initial recruiter phone screen through technical rounds and final leadership interviews. Maintaining prompt communication with your recruiting coordinator helps keep the process moving efficiently.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.