As a Data Scientist at Accenture Federal Services, you play a pivotal role in helping the United States federal government solve complex operational, strategic, and security challenges. You bridge the gap between raw, multi-source data and actionable intelligence, designing analytical workflows that directly influence defense, national security, public safety, and civilian agency missions. Your day-to-C day work involves translating ambiguous mission requirements into robust data pipelines, machine learning models, and intuitive dashboards that empower stakeholders to make rapid, data-driven decisions.
This role requires balancing technical execution with high-touch stakeholder collaboration across diverse federal client environments. You will work alongside data engineers, product managers, and domain specialists to extract insights from massive, secure datasets while maintaining strict data governance. Whether you are deploying advanced natural language processing models, fine-tuning large language models, or architecting custom analytics on cloud platforms like AWS and Databricks, your contributions directly impact the safety, efficiency, and effectiveness of government programs.
Succeeding in this role demands a unique combination of core technical competence and adaptability. You must be comfortable navigating secure infrastructure, writing optimized code in Python and SQL, and communicating complex technical concepts to non-technical leaders. Expect an environment where intellectual curiosity, rigorous quantitative analysis, and a commitment to public service are deeply valued and rewarded.
Recruiter Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Screening
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Core Interview Rounds
reportedRounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.
What to demonstrate
- Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
- The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
- Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options
How to prepare
- Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
- Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
- Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
Behavioral Interviews
reportedMost of the weight in this round sits on the disagreement questions. Data work routinely produces an answer someone senior did not want, and the interviewer is trying to learn what you do in that hour. Both failure modes are common: folding as soon as a director pushes back, and treating the pushback as ignorance to be corrected with a better chart. A strong answer usually contains a specific thing the other person knew that you did not, and describes how you found out whether it changed the conclusion.
What to demonstrate
- Whether you can state the other side's argument accurately before you explain why you disagreed
- What you treated as evidence during the disagreement, such as a rerun under their assumption or a holdout check, rather than persuasion technique
- Whether you distinguish being overruled from being wrong, and can give an example of each
How to prepare
- Write out one disagreement where you turned out to be wrong, and say what in the data misled you. Candidates prepare the story where they were right, and the follow-up asks for the other one.
- For your main disagreement story, be ready to say what result would have made you drop your position. If no such result exists, you were not arguing from the data.
- Practise stating the opposing position out loud in one sentence the stakeholder would accept, then continue the story.
Technical Case Studies
reportedUnderneath the business framing, this round is usually asking whether you can turn a fuzzy goal into a quantity that could be computed from data such a business would plausibly hold. That means a metric with a stated numerator, denominator, eligibility rule and time window, plus an honest account of the conditions under which it would mislead you. Answers come apart when a candidate names a familiar metric and never defines it, because every follow-up then lands on an ambiguity that was left open and the candidate has to invent the definition under pressure.
What to demonstrate
- Whether a named metric arrives with its denominator, eligibility rule and window attached rather than assumed
- Whether the measure follows from the mechanism you proposed, or is a recognisable metric retrofitted to it afterwards
- Whether you name a guardrail that would reveal the gain came from somewhere you did not want it to come from
- Whether you can say what data the plan requires and what you would settle for if that logging were never implemented
How to prepare
- Take five metrics you reach for by reflex and write each as one sentence containing numerator, denominator, eligibility rule and time window. The ones you cannot finish are the ones that will fail under follow-up.
- For a product you use daily, write the measurement plan you would propose for a change to it: primary metric, one guardrail, the unit of analysis, and the table the numbers would come from.
- Practise the substitution question. For three metrics you like, write what you would measure instead if the event you depend on were not being logged.
PracHub editorial advice for the preparation topics above.
Computing coverage or take-up with an administrative denominator, for example approved applications divided by all applications
That ratio measures throughput among people who already found the system, not delivery to the people entitled to the service. It gets better when outreach is cut, because the marginal applicant is the one most likely to be denied or to abandon, and it gets worse when a new access channel brings in harder cases. The correct denominator is a modelled eligible population from survey microdata run through the eligibility rules, and the gap between it and the applicant count is usually the finding.
Dividing an administrative count by a survey population estimate and reporting the rate as if it were exact
At block-group and tract scale the published margin of error is often a large fraction of the estimate, so the variance of the resulting rate is dominated by the denominator rather than the numerator, and a leaderboard of areas ranked by point estimate is largely a leaderboard of the smallest and noisiest areas. Propagate the denominator error into the rate, or aggregate up until the relative standard error is acceptable, and state the threshold used. The same join also breaks silently across boundary vintages: a geo_code can refer to a different physical area after a redraw, so joining current-vintage geography onto historical events reassigns records and manufactures a trend break.
Generalising beyond the population the sample actually supports
State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.
Analysing at a different unit than the one randomised
Say out loud what was randomised (user, device, account, cluster) and make the analysis unit match, or account for the clustering with cluster-robust standard errors, the delta method, or aggregation up to the randomised unit. Randomising users and then running a test over sessions understates variance and inflates the false-positive rate.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the false discovery rate and how you adjust p-values when cond…
Explain the false discovery rate and how you adjust p-values when conducting thousands of simultaneous anomaly detection checks.
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
- Translate the result into the decision it informs, in one plain sentence.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
How do you test for homoscedasticity and normality in residuals when b…
How do you test for homoscedasticity and normality in residuals when building a predictive regression model for resource allocation?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Say how the offline result would be validated online before it is trusted.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Publish on-time decision share anchored on the due date
You are given applications as a DataFrame with application_id, program_code, submitted_at, sla_due_at, decision_at (NaT while open) and status, plus a scalar extract_date. Implement on-time decision share monthly, keyed on the month of sla_due_at: the denominator is every application whose sla_due_at falls in that month whether or not it has been decided, and the numerator is those with decision_at <= sla_due_at. Flag months that are not yet complete. Then report, for one recent month, the value you would have published had you keyed on decision_at over decided rows only.
Approach
- Restrict to rows with submitted_at populated. Drafts and abandoned applications never acquire an sla_due_at, and sweeping them in gives a denominator that moves with intake abandonment rather than with adjudication.
- Build the denominator by flooring sla_due_at to month and counting every row in that month, including rows where decision_at is NaT. Keeping undecided rows in the denominator is the entire reason the definition anchors on the due date, so this is the line to write first.
- Numerator is decision_at.notna() & (decision_at <= sla_due_at). NaT comparisons already evaluate False, but write the notna() explicitly so the intent survives someone later swapping in a fillna.
- Mark a month incomplete when its due dates extend past extract_date, since a case whose due date is still in the future can yet be decided on time. Publishing those as provisional stops the current month reading as a collapse.
- Recompute keyed on decision_at over decided rows only and publish it beside the primary series rather than subtracting one from the other. The two denominators are different sets, every application due in the month against every application decided in the month, so the month-level gap between the two shares is not a population and its sign is not predictable. Report open_past_due as its own column instead, since that count is what the decision-keyed version cannot see.
Worked solution 20 min
- Filter to submitted_at.notna() and sla_due_at.notna(), counting any submitted row that is missing a due date rather than letting it drop silently, and derive due_month = sla_due_at.dt.to_period('M').
- Group by due_month: denominator is the row count, numerator is the sum of (decision_at.notna() & (decision_at <= sla_due_at)).
- Add is_complete by comparing the month end against extract_date, and carry a column for open_past_due = decision_at.isna() & (sla_due_at < extract_date).
- Separately group decided rows by decision_at month and compute the same ratio over that base for the comparison column.
- Return one frame with both series and the open past-due count, so the gap between them is readable without a second query.
Follow-up
- A unit spends one week clearing 400 cases that were already past their due date. The decision-keyed share for that month drops hard and the due-date-keyed series does not move in any month. Explain both to a director in two sentences and say which number you would publish.
- Decision events in the case log can be backdated by staff. Which clock should this metric use, and how would you measure whether the choice changes the published number?
- The service standard itself changed mid-month under a new rule version, so sla_due_at is derived from two different standards inside one bucket. How do you report that month?
Given a table of user access logs, write a query to find the top three…
Given a table of user access logs, write a query to find the top three most frequent activities per department using ranking functions.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a query using SQL window functions to calculate a running seven-…
Write a query using SQL window functions to calculate a running seven-day moving average of transaction volumes across multiple regional databases.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Reconcile backdated events against their recording timestamps
fact_case_event stores event_at and recorded_at as TIMESTAMPTZ in UTC: event_at is when staff assert the action happened, recorded_at is when the row was written, and is_backdated marks event_at falling before the date of recorded_at. The agency's business day is defined by a local timezone name supplied as a parameter. For one quarter and event_type = 'decision', return a daily series carrying decisions counted on the local event_at date, decisions counted on the local recorded_at date, the median entry lag in hours, and a flag where the two counts differ by more than 20 percent.
Approach
- Convert before truncating: (event_at AT TIME ZONE :agency_tz)::date yields the local wall-clock date, whereas event_at::date silently uses the session timezone and pushes every evening decision onto the following UTC day.
- Build the two counts on separate day keys and FULL OUTER JOIN them on the date, or aggregate from a generated date spine, because a day can have recordings with no assertions or the reverse and an inner join would drop exactly the interesting days.
- Compute the lag as EXTRACT(EPOCH FROM (recorded_at - event_at)) / 3600.0 and take percentile_cont(0.5) WITHIN GROUP (ORDER BY lag_hours); the median is the right summary because the lag distribution has a long right tail from bulk backdating.
- Key the lag to one of the two day definitions and say which, since the median lag on the recorded_at day answers a different question from the median lag on the event_at day.
- Express the divergence flag as ABS(a - b) > 0.20 * GREATEST(a, b) with a guard against a zero denominator, and expect it to cluster at fiscal-period close rather than scatter randomly.
Worked solution 30 min
- Write both day keys for a handful of rows side by side and confirm that rows whose local time is late in the evening get different dates under the two casts.
- Aggregate decisions by local event_at date into one CTE and by local recorded_at date into another.
- FULL OUTER JOIN the two on date and COALESCE the counts to zero.
- Add the median lag from a third aggregate keyed on the local recorded_at date and join it in.
- Compute the divergence flag and inspect where the flagged days fall in the quarter.
Follow-up
- An interrupted time series keyed on event_at shows a clean level shift on the policy effective date. What would you check before believing it?
- Which of the two clocks should the published on-time decision share use, and what would you write in the footnote?
- Half the quarter's backdated rows land in the final three days. Is that fraud, a batch job, or normal practice, and how would you tell them apart from these columns alone?
How do you evaluate the trade-off between Type I and Type II errors in…
How do you evaluate the trade-off between Type I and Type II errors in a high-stakes fraud detection system?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How would you define the core metrics for a newly launched search tool…
How would you define the core metrics for a newly launched search tool designed to help case workers quickly surface relevant historical records?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design a product metric framework to measure the operati…
How would you design a product metric framework to measure the operational success of an investigative dashboard used by federal law enforcement agencies?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Give an example of how you managed competing priorities across multipl…
Give an example of how you managed competing priorities across multiple cross-functional teams while adhering to a strict project milestone.
Approach
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What are the most common experimentation pitfalls, such as sample rati…
What are the most common experimentation pitfalls, such as sample ratio mismatch or premature stopping, and how do you prevent them?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- What would you do if you could not randomise at all?
Explain how you ensure statistical significance when running concurren…
Explain how you ensure statistical significance when running concurrent tests across overlapping user segments in a government web application.
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
What framework would you use to evaluate whether an AI-driven recommen…
What framework would you use to evaluate whether an AI-driven recommendation feature actually improves user decision-making speed?
Approach
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Evaluate a staggered channel rollout without a control arm
An online renewal channel went live in 9 of 46 service districts on six different dates across 14 months, ordered by district readiness. There is no holdout. Using fact_application and fact_case_event, estimate the effect on the on-time decision share, which is anchored on sla_due_at rather than decision_at. Name the estimator, state the identifying assumption and how you would probe it, and say what you would do differently if only one district had adopted.
Approach
- Refuse the default: two-way fixed effects under staggered timing with effects that change over time forms comparisons that use already-treated districts as controls, and its single coefficient is a variance-weighted average that can carry negative weights, so it is not the average effect anyone is asking for.
- Estimate group-time effects directly against a not-yet-treated comparison group, then aggregate the ATT(g,t) into an event-study path and one post-adoption summary, so every comparison is a newly treated district against a district that has not yet adopted.
- Take the assumption seriously given how adoption was ordered: readiness-ordered rollout means early districts differ from late ones in the dimension the outcome responds to, so plot pre-treatment leads and then run a sensitivity analysis over bounded violations of parallel trends, because a non-significant pre-trend test at 9 treated units proves very little.
- Fix the clock before fitting anything: define adoption on the policy go-live date, reconcile event_at against recorded_at for backdated and batch-written rows, and confirm the outcome window keys on sla_due_at so a case still undecided past its due date counts as a failure instead of leaving the sample.
- Get inference right at 46 clusters with 9 treated: cluster-robust standard errors are unreliable at that treated count, so use a wild cluster bootstrap or randomisation inference over the possible adoption schedules.
- If adoption were a single district, switch to synthetic control: fit donor weights on a long pre-period of the same outcome plus predictors, require a good pre-period fit before making any claim, and draw inference from placebo permutations across donors using the post-to-pre RMSPE ratio.
Worked solution 45 min
- Build a district by month panel of the on-time share with the denominator anchored on sla_due_at, and mark each district's adoption month g.
- Estimate ATT(g,t) using not-yet-treated districts as the comparison, then aggregate to an event-time path from -6 to +12 and to an overall post-adoption summary.
- Plot the leads, and separately report the smallest violation of parallel trends that would overturn the sign of the estimate.
- Re-run inference with a wild cluster bootstrap at the district level.
- As a robustness check, drop the two earliest adopters, which contribute the longest post window and the least comparable comparison set.
Follow-up
- The event study shows a downward lead two periods before adoption. Name three explanations and how you would tell them apart.
- Districts that adopted later had worse backlogs. Does that break parallel trends, and can any of your estimators survive it?
- How would you detect a queue flush just before go-live to clear the backlog, and what would that do to your estimate?
Weekly submissions drop the week a fee takes effect
Submitted applications for one program fell 22 percent week over week, the same week a fee increase took effect. You have fact_application (application_id, program_code, channel, created_at, submitted_at, status) and fact_case_event (application_id, event_type, event_at, recorded_at, actor_type). Leadership wants a same-day read on whether demand actually fell. Deliver a determination with the ordered checks you ran, and if the drop is an artefact, the corrected week-over-week figure.
Approach
- Count the week on both clocks before interpreting anything: applications by date(submitted_at), and the matching fact_case_event rows of event_type = 'created' by date(recorded_at). A real demand change moves both series together; a load or ingestion failure moves recorded_at only, leaving a flat-zero gap on specific load dates while submitted_at values stay spread across the prior days they claim to belong to.
- Cut by channel. Intake arrives through online, phone, mail, in_person, kiosk and partner_org, and each has a different upstream. A drop concentrated in one channel at roughly 100 percent of that channel's volume is a feed, not a behaviour; a fee response should appear across the channels that expose the fee and be graded, not total.
- Check freshness per channel with MAX(recorded_at) and a row count by load date. Mail and partner_org batches land days late by design, so the most recent week is structurally incomplete on those channels and must not be compared against a matured week.
- Separate the funnel stages. If created_at volume is flat and only submitted_at fell, constituents are starting and abandoning, which is the shape a fee at the submit step would produce. If both fell together, the change is upstream of the form.
- Align the comparison windows on weekday and on published closure days before quoting any percentage, then re-run the identical query after the late batches land and report the restatement alongside the original.
Follow-up
- The partner_org batch lands three days later and the gap closes. What monitor would have caught this before leadership saw the number, and what is its alert condition?
- Suppose created_at is flat and submitted_at genuinely fell 9 percent. How would you attribute that to the fee rather than to the seasonal pattern in the same weeks of prior years?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Have two ready. In one, the data was on your side and you had to move someone who outranked you. In the other, the pushback was correct and you changed position. The second is the harder story and it lands better, because it shows you separate being right from being attached to an answer. Name the person's actual objection.
How do you handle network effects and interference between treatment a…
How do you handle network effects and interference between treatment and control groups when testing a collaborative software feature?
Approach
- Quantify the outcome, including what you would not claim credit for.
- State the situation in two sentences and spend the rest on your reasoning.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know the outcome was caused by your change?
Disagree with a proposal to auto-deny on unreturned evidence
A product manager proposes auto-denying any application sitting in status 'pending_evidence' for 14 days, writing denial_reason_code = 'evidence_not_returned', and projects a large fall in median days to decision. You have fact_application, fact_case_event and dim_geography. Write a one-page disagreement: the three numbers you would compute before the change ships, what each would have to show for the proposal to be sound, and the result that would make you drop your objection.
Approach
- Name the mechanism plainly: a procedural denial removes a case from the pending pool without a substantive decision, so the timeliness improvement is arithmetic. Any argument that does not address this is a disagreement about values rather than about the proposal.
- Number one, appeal overturn rate split by procedural versus substantive denial_reason_code, on a denial cohort with a 180-day follow-up. If procedural denials are overturned or remanded far more often, the rule manufactures rework that lands in cost per sustained outcome rather than in the timeliness number.
- Number two, the distribution of the lag from evidence_requested to evidence_received in fact_case_event, split by channel and preferred_language_code. If the mass beyond 14 days is concentrated in mail and in-person intake or in non-English speakers, the threshold has a disparate impact before any model is involved, and comparing like with like means holding evidence_item_code fixed.
- Number three, the reapplication rate: how often an evidence_not_returned denial is followed by a new fact_application row from the same constituent_id within 90 days. Those are the same case counted twice and the throughput gain is illusory.
- State the falsifier explicitly. Flat overturn rates across codes, a thin tail past 14 days that does not vary by channel or language, and rare reapplication would make the rule reasonable, and you will say so in writing.
- Propose the cheaper decision: a staged rollout by adjudicator_unit_id with pre-registered guardrails and a stopping rule, rather than a blanket switch that cannot be evaluated afterwards.
Follow-up
- The product manager says the published service standard requires the case to close. Does that change your position, and what do you ask for instead?
- Design the staged rollout: which units, which comparison, and what result stops it.
Explain why a ranked district table should not be published
You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.
Approach
- Convert each published margin to a standard error first: at 90 percent confidence the standard error is moe divided by 1.645. Compute the relative standard error of the denominator for every area in the table.
- Show where the uncertainty lives. Treating the administrative count as fixed and independent of the survey estimate, the rate's relative standard error equals the denominator's, so an area with a 24 percent relative standard error on population has a 24 percent relative standard error on its rate no matter how clean fact_service_request is.
- Demonstrate the instability rather than asserting it: resample each denominator from its published standard error, recompute the ranking a few thousand times, and report how often each area actually lands in the worst twelve. Areas that appear in only a third of draws are not findings.
- Replace the rank with something that survives the uncertainty: aggregate to a geo_level where the relative standard error clears a threshold you state out loud, or group areas into tiers whose intervals do not overlap, and name the threshold as a choice you made rather than a standard.
- Give the executive one sentence she can repeat without you in the room, for example that the data supports naming a group of high-demand areas but not ordering them.
Follow-up
- Which relative standard error threshold do you use, and why is it a judgement rather than a rule?
- Two of the areas you aggregated sit across a boundary redraw. What breaks in the join, and how would you notice?
- 01
How do you handle network effects and interference between treatment and control groups when testing a collaborative software feature?
- 02
A product manager proposes auto-denying any application sitting in status 'pending_evidence' for 14 days, writing denial_reason_code = 'evidence_not_returned', and projects a large fall in median days to decision. You have fact_application, fact_case_event and dim_geography. Write a one-page disagreement: the three numbers you would compute before the change ships, what each would have to show for the proposal to be sound, and the result that would make you drop your objection.
- 03
You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.
Is this an official Accenture Federal Services interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Accenture Federal Services. Rounds and questions reflect what candidates have reported, not a process Accenture Federal Services has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What level of interview difficulty should I expect for the Data Scientist role?
The interview loops range from straightforward conversational rounds to rigorous technical assessments involving live coding, system design, and statistical problem-solving. Preparation should cover both foundational coding in Python and SQL as well as high-level product and experimental design principles.
PracHub interview research ↗How important is an active security clearance for getting hired?
An active security clearance (such as Secret, Top Secret, or TS/SCI with polygraph) is often a mandatory prerequisite depending on the specific client engagement and team. Be sure to clarify clearance requirements with your recruiter early in the screening process.
PracHub interview research ↗How long does the typical interview process take from initial screen to offer?
The timeline varies based on scheduling availability and clearance verification, but typically spans several weeks from the initial recruiter phone screen through technical rounds and final leadership interviews. Maintaining prompt communication with your recruiting coordinator helps keep the process moving efficiently.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22