The Institute for Defense Analyses (IDA) is a non-profit corporation that operates three Federally Funded Research and Development Centers (FFRDCs) for the United States government. As a Data Scientist at the Institute for Defense Analyses, you do not work on optimizing ad click-through rates or maximizing social media engagement. Instead, your work directly impacts national security, military readiness, and high-stakes government policy. You will apply advanced statistical modeling, machine learning, and operations research to solve complex, unstructured problems for the Department of Defense and other federal agencies.
Because the Institute for Defense Analyses prides itself on providing objective, scientifically rigorous, and conflict-of-interest-free analyses, the role of a Data Scientist here demands an exceptionally high standard of intellectual honesty. You will work alongside multidisciplinary teams of physicists, economists, engineers, and military experts to evaluate defense systems, assess cyber security vulnerabilities, and optimize resource allocation. The insights you generate will frequently be delivered directly to senior leadership within the Pentagon and other executive-level decision-makers.
Securing a Data Scientist position at the Institute for Defense Analyses requires demonstrating not only top-tier technical proficiency but also the ability to defend your methodology under intense scrutiny. The work culture mirrors the academic rigor of a top-tier research university combined with the mission-driven focus of the national security sector. If you are motivated by intellectually demanding problems that have a tangible impact on global stability and national defense, this role offers an unparalleled platform.
Application Submission
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
Virtual Screening
reportedAn extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.
What to demonstrate
- Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
- Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
- How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one
How to prepare
- Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
- Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
- Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
Onsite Evaluation
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub editorial advice for the preparation topics above.
Treating event_at in the case event log as when the event actually happened
Staff backdate to the true date of an action and enter it later, and batch loads write many rows at once, so event_at and recorded_at diverge systematically and the divergence is largest at fiscal-period and reporting-deadline boundaries. An interrupted time series keyed on event_at will read a queue flush before a deadline as a level shift caused by the policy change, and a dashboard keyed on recorded_at will show activity on days when nothing happened. Reconcile both columns, model the entry lag distribution, and state which clock each metric uses.
Computing coverage or take-up with an administrative denominator, for example approved applications divided by all applications
That ratio measures throughput among people who already found the system, not delivery to the people entitled to the service. It gets better when outreach is cut, because the marginal applicant is the one most likely to be denied or to abandon, and it gets worse when a new access channel brings in harder cases. The correct denominator is a modelled eligible population from survey microdata run through the eligibility rules, and the gap between it and the applicant count is usually the finding.
Crediting a treatment for regression to the mean
Selecting a group because it is extreme (lowest-engagement users, accounts having their worst month, the bottom decile of a score) moves that group's expected next-period value back toward the average even under no treatment, by exactly as much as the selecting measure is imperfectly correlated with its own later value. Compare against units that met the same selection rule and went untreated, or use two pre-periods so the bounce-back is visible before the intervention starts. A pre-post number on a group chosen for being extreme measures the selection rule, not the treatment.
Ending an analysis without a recommendation or next step
Close with what you would do and what would change your mind, stated as a condition you can check later. If the evidence is genuinely inconclusive, recommend the specific next measurement and say what it costs in time or exposure.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the difference between generative and discriminative models, a…
Explain the difference between generative and discriminative models, and give an example of when you would use each in a predictive threat-assessment scenario.
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- Which assumption here is most likely to be violated in practice?
- What sample size would you need to detect an effect half this size?
How do you detect and handle multicollinearity in a dataset with hundr…
How do you detect and handle multicollinearity in a dataset with hundreds of potential policy indicators?
Approach
- Write down the assumption the method needs before you use the method.
- Translate the result into the decision it informs, in one plain sentence.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
What is the curse of dimensionality, and what techniques do you use to…
What is the curse of dimensionality, and what techniques do you use to mitigate it when working with highly complex sensor data?
Approach
- Write down the assumption the method needs before you use the method.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
Explain a scenario where your data violated the assumptions of your st…
Explain a scenario where your data violated the assumptions of your statistical model. How did you adjust your approach?
Approach
- Say how the offline result would be validated online before it is trusted.
- Check what information would not exist at prediction time, and exclude it.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
How do you balance model interpretability with predictive power when p…
How do you balance model interpretability with predictive power when presenting findings to non-technical defense sponsors?
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Laplace release of nested counts with consistent post-processing
counts holds exact denial counts by (district_geo_id, tract_geo_id, denial_reason_code), where tracts nest inside districts and each constituent contributes to exactly one cell. Release the tract by reason cells and the district totals under a total privacy budget epsilon, using a Laplace mechanism you write yourself. Then post-process each district so its tract cells are non-negative and sum exactly to the released district total. Finally show empirically how per-query error moves when one epsilon is split across k sequential queries.
Approach
- Fix the adjacency and the sensitivity before touching data. The full tract by reason histogram has L1 sensitivity 1 under add-or-remove adjacency and 2 under replace-one, because a replaced record leaves one cell and enters another. Say which you are using; it doubles the noise scale b = sensitivity / epsilon_cells.
- Split the budget honestly. District totals are a coarsening of the same records rather than a disjoint query, so releasing them alongside the cells composes sequentially and epsilon_cells + epsilon_totals = epsilon. Disjointness buys you something only across records, for example separate districts released by separate mechanisms.
- Add noise with a seeded generator: rng.laplace(0, b, size). Laplace(b) has variance 2b squared, so per-cell standard deviation is sqrt(2) * b. Compare that number against the typical cell count before deciding the release is publishable at all.
- Post-process per district: clip the noisy district total at 0, then project the noisy tract vector onto {x >= 0, sum x = T} by minimising squared distance. The solution is x_i = max(y_i - tau, 0) with tau found by sorting y descending and scanning for the value that makes the sum hit T. If integers are required, round with largest remainders so the total survives the rounding.
- Run the sweep: for k in 1 to 8, split one epsilon into k sequential queries each with b = k * sensitivity / epsilon, and plot RMSE against k. Per-query standard deviation grows linearly in k, which is the whole argument for releasing fewer and coarser queries.
- State why the reconciliation is free: any function of a differentially private output is differentially private with the same parameters, provided it touches no raw data again.
Worked solution 45 min
- Write laplace_release(counts, sensitivity, epsilon, rng) returning noisy values, and use it once for the tract by reason cells and once for the district totals.
- Implement project_to_simplex(y, T) with the sorting-and-threshold search, and assert its two constraints on the way out.
- Apply the projection per district, then largest-remainder rounding if integers are required, and assemble the released frame.
- Sweep k from 1 to 8 with b = k * sensitivity / epsilon, recording RMSE against the true counts over several seeds per k.
- Report per-cell noise standard deviation, the RMSE curve, and the smallest aggregation level at which noise standard deviation falls under 5 percent of the cell count.
Follow-up
- You could skip epsilon_totals and derive district totals by summing the noisy tract cells. Compare the variance of that against a directly measured total and say when each wins.
- A reason code appears in only two tracts, with a true count of 3 in each. What do you publish, and does the noise alone make it safe?
- The same table is released monthly for a year. What is the honest statement about cumulative budget, and what would you change in the design?
Total time each case spent awaiting evidence
fact_case_event is append-only with case_event_id, application_id, event_type, event_at, prior_status and new_status. A case enters the pending_evidence state on an event where new_status = 'pending_evidence' and leaves it at that application's next event in event_at order; it may enter more than once, and it may still be sitting there at the extract timestamp. Return per application_id the number of pending_evidence stints and the total hours spent in that status, counting any open stint up to the extract timestamp. Order events within an application by event_at with case_event_id as the tiebreaker.
Approach
- Compute LEAD(event_at) OVER (PARTITION BY application_id ORDER BY event_at, case_event_id) over the full event stream in a CTE, before filtering to pending_evidence rows: if you filter first, LEAD returns the next entry into pending_evidence rather than the exit from it, and every stint is measured from one entry to the following entry.
- Include case_event_id in the ORDER BY because staff backdating puts several events on the same event_at, and a window with a non-deterministic order silently reorders a transition pair.
- In the outer query, keep only rows where new_status = 'pending_evidence' and set the stint end to COALESCE(next_event_at, extract_ts), which is what makes open stints count instead of vanishing.
- Aggregate with COUNT(*) for stints and SUM(EXTRACT(EPOCH FROM (stint_end - event_at)) / 3600.0) for hours, grouping by application_id.
- Flag applications whose final stint is open separately, because those durations are right-censored lower bounds and must not be pooled into a mean as if they were complete.
Worked solution 30 min
- Build a CTE selecting all fact_case_event columns plus LEAD(event_at) OVER (PARTITION BY application_id ORDER BY event_at, case_event_id) AS next_event_at.
- Verify on one multi-stint application by listing its events with next_event_at alongside, and confirm each pending_evidence row's next_event_at is an evidence_received or decision event.
- Filter the CTE to new_status = 'pending_evidence' and compute stint_end = COALESCE(next_event_at, extract_ts).
- Aggregate stints and hours per application_id, and carry a BOOL_OR(next_event_at IS NULL) flag for open stints.
- Compare the total hours on the same application computed with and without the COALESCE to size how much the open stints contribute.
Follow-up
- The average of this total is used to claim evidence handling improved. Why is that average biased downward while the backlog is growing, and what would you report instead?
- Two events on one application share an event_at and describe opposite transitions. How do you decide which came first?
- How would you extend this to give the time in every status rather than just pending_evidence, in one pass?
Award value from signed actions, ceilings and outlays
fact_obligation holds one row per funding action: obligation_action_id, award_id, action_type, obligated_delta_cents (signed, negative for deobligation and termination), ceiling_amount_cents, outlay_to_date_cents, action_at, fiscal_year, snapshot_date, is_current_snapshot. Within a single snapshot, ceiling_amount_cents and outlay_to_date_cents are award-level values repeated on every action row of that award; obligated_delta_cents is per action. Return one row per award_id active in fiscal_year 2026 with net obligated, current ceiling, outlay to date, and outlay divided by net obligated. Drop awards whose net obligated is not positive.
Approach
- Filter is_current_snapshot = TRUE before any aggregation; the table is a stack of as-of views, so an unfiltered SUM adds every historical snapshot of the same award together and inflates the total by roughly the number of snapshots retained.
- Select the award population with a semi-join on fiscal_year: inside the current snapshot take the DISTINCT award_id values having at least one action row with fiscal_year = 2026, then aggregate against that set. Skipping this step is not a narrower answer, it is a different one — the query returns every award in the snapshot, including awards whose last action was years earlier.
- Do not push fiscal_year = 2026 into the aggregation itself. That sums only fiscal-2026 deltas while ceiling_amount_cents and outlay_to_date_cents stay award-to-date values, so the ratio divides lifetime cash by one year of obligations and runs above 1.0 on every multi-year award. Scope is per award; net obligated, ceiling and outlay are all award-to-date.
- Sum obligated_delta_cents, which is the only column where summation is correct, because a modification or deobligation is an increment against the award and not a restatement of it.
- Take ceiling_amount_cents and outlay_to_date_cents with MAX (or from the latest action_at) rather than SUM: they are repeated award-level facts within the snapshot, so summing multiplies them by the action count.
- Guard the ratio with NULLIF on the denominator, then read a ratio above 1.0 as a finding rather than a number to report, since cash out cannot exceed money committed under normal accounting.
- State in the output or the header that ceiling is not money committed, so nobody downstream sums the ceiling column as a spending figure.
Follow-up
- An award has actions in both fiscal 2025 and 2026. What does 'net obligated for fiscal 2026' mean, and which of the two plausible definitions would you ship?
- Outlay-to-obligation ratios cluster near zero for awards signed in the last quarter. Is that a data problem?
- How would you detect an award that was terminated, given only action_type and the signed delta?
What metrics did you use to evaluate the performance of your model, an…
What metrics did you use to evaluate the performance of your model, and why were those metrics appropriate for that specific problem space?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you design an experiment to test the effectiveness of a new …
How would you design an experiment to test the effectiveness of a new military training simulation when sample sizes are extremely limited?
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
How did you handle missing data or high levels of uncertainty in your …
How did you handle missing data or high levels of uncertainty in your previous research? What assumptions did you make, and how did you validate them?
Approach
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Separate obligations, outlays and ceilings in a funding execution metric
fact_obligation holds one row per funding action with obligation_action_id, award_id, action_type, obligated_delta_cents (signed), ceiling_amount_cents, outlay_to_date_cents, action_at, fiscal_year, fiscal_quarter, program_code, competition_type, period_of_performance_start, snapshot_date and is_current_snapshot. A proposed execution rate divides obligations by ceiling. Define the execution metric you would report instead, write the aggregation rule that avoids double counting across snapshots and across actions within an award, and give the guardrails that catch obligating money late in the fiscal year to avoid a lapse rather than to deliver anything.
Approach
- Reject the proposed ratio on a definitional ground: ceiling_amount_cents is the maximum an award may reach, not money committed, so obligations over ceiling measures how much of an optional headroom has been used and rises whenever ceilings are set tightly. It is not an execution rate.
- Define execution as outlays against obligations on a fixed obligation vintage. Take awards whose obligating actions fall in fiscal year t, then measure cumulative outlay at t plus 12 and t plus 24 months. Cash lags obligation by months to years, so an execution rate measured inside the same fiscal year is mechanically near zero and says nothing about delivery.
- Write the aggregation rule in two parts because the table has two distinct double-counting traps. First, the table is a series of as-of views, so an unfiltered query sums the same action once for every snapshot that contains it: filter to one snapshot_date or is_current_snapshot = TRUE. Note that this inflation factor is not uniform, since an action is absent from every snapshot taken before it was recorded, so it cannot be divided out after the fact. Second, obligated_delta_cents is per action and must be summed across actions within award_id, with deobligations and terminations entering as negatives, while outlay_to_date_cents is a cumulative as-of value carried on the award and repeats across that award's action rows, so summing it across actions multiplies it by the action count.
- Add the fiscal-timing guardrail: share of the fiscal year's obligated_delta_cents with action_at in the final 30 days, compared against the 1 in 12 share that an even distribution would give. A concentration there is the signature of obligating to avoid a lapse rather than to deliver.
- Add two guardrails that catch what year-end haste costs. Deobligation and termination rate within 12 months of the original award, which catches money committed to work that was never viable, and the competition_type mix, because the fastest way to obligate quickly is to move volume from full_and_open toward sole_source.
Worked solution 40 min
- Write the award-level query: filter to is_current_snapshot = TRUE, group by award_id, SUM(obligated_delta_cents) for award value, and take outlay_to_date_cents from a single row per award rather than summing it.
- Assign each award to an obligation vintage by the fiscal_year of its first new_award action, and hold that assignment fixed across later evaluations.
- Compute execution at 12 and 24 months after vintage close and report both, since a single horizon cannot distinguish slow start from failure to deliver.
- Compute the year-end concentration guardrail and compare it against the even-distribution benchmark.
- Compute 12-month deobligation rate and competition_type mix on the same vintage, and place all three guardrails beside the headline.
Follow-up
- An award is modified upward in fiscal year t plus 1. Which vintage does the new money belong to for the execution metric, and what does your choice do to the denominator of the original vintage?
- Outlay data itself arrives with reporting lag. How would you tell a genuine execution problem apart from a reporting lag in the outlay feed?
Quarterly obligations double in an as-of snapshot table
Obligated dollars for one program appear to double in a fiscal quarter. fact_obligation holds one row per funding action per snapshot: obligation_action_id, award_id, action_type, obligated_delta_cents, ceiling_amount_cents, outlay_to_date_cents, fiscal_year, fiscal_quarter, snapshot_date, is_current_snapshot. Establish whether obligations actually rose, produce the corrected quarterly series, and name the comparison period you would publish it against.
Approach
- Check snapshot hygiene before any aggregation. Count distinct snapshot_date per obligation_action_id: if actions repeat across snapshots, an unfiltered SUM counts the same action once per snapshot. Restrict to is_current_snapshot = TRUE, or to a single chosen snapshot_date, and re-run.
- Sum the signed obligated_delta_cents across all action types rather than filtering to action_type = 'new_award'. Deobligations and terminations are negative rows, so dropping them inflates every period; award value is only correct when summed across all actions grouped by award_id.
- Confirm no one has substituted a different money column. ceiling_amount_cents is a maximum the award may reach and is not money committed; outlay_to_date_cents is cumulative cash as of the snapshot and lags obligation by months to years, so differencing it across snapshots gives a flow while summing it gives nonsense.
- Choose the comparison period on the basis of how the appropriation behaves. Where funds lapse at fiscal year end, obligations concentrate in the closing weeks, so quarter over quarter compares a peak against a trough; compare the same fiscal quarter across years instead, and say so in the note.
- If the rise survives all of that, decompose by award_id, action_type and competition_type. A single option exercise or one multi-year award booked in full is a different finding from broad growth, and the two should never be reported with the same sentence.
Follow-up
- Obligations rose and outlays did not. What are the benign explanations, and which one would you test first?
- How would you publish a series that is stable when a prior quarter is restated in a later snapshot?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.
How do you handle receiving intense, highly critical feedback on your …
How do you handle receiving intense, highly critical feedback on your work from peers or senior researchers?
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
- Quantify the outcome, including what you would not claim credit for.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Turn "how many people are we missing" into a scoped estimate
A programme director stops you and asks, "how many people are we missing?" You have fact_application, dim_constituent and dim_geography and nothing else. Before writing any SQL you owe her a written scope in thirty minutes: the question restated as something estimable, the denominator you will build and where it comes from, what the data cannot tell her, and a delivery date. Produce that scope note in under 300 words, plus the three clarifying questions whose answers would change the estimate.
Approach
- Restate the request as a coverage estimand before negotiating anything else: distinct constituent_id with status = 'approved' and decision_at inside a named program year, over a modelled count of eligible constituents in the same geography and year. Say out loud that the denominator cannot come from fact_application, because every row there belongs to someone who already found the system.
- Write the three questions whose answers move the number by more than the analysis will: which program_code and which rule_version_id define eligibility, which geo_level and vintage_year is the denominator drawn at, and whether "missing" means never created an application, created one but has submitted_at IS NULL, or was denied on a procedural denial_reason_code. Those are three different interventions and only the first is a pure outreach problem.
- Name the denominator pipeline concretely: survey microdata plus a microsimulation of the eligibility rule in force at that rule_version_id, with the income and household definitions stated because reported_income_cents in the administrative table is self-reported and heaps on round values.
- State the non-answers before they are asked for: the estimate is a count, not a list of people, and the gap cannot be attributed to a cause without a design.
- Commit to a date and offer an interim artefact on day one, the funnel from created_at to submitted_at to decision_at by channel, so the director has something concrete while the denominator is built.
Follow-up
- The director asks for the names of the people who are missing so outreach can call them. What do you say, and what could you legitimately produce instead?
- Your modelled eligible population comes out 40 percent above the applicant count. What checks do you run before anyone sees that number?
Allocate one analyst-week across three competing requests
Three requests land on Monday and you have one analyst-week. A statutory report of decision counts by program_code is due in nine days. An operations team wants a backlog projection to size next quarter's staffing. An equity lead wants the first-response equity ratio by deprivation quartile for a briefing in six weeks. Produce the allocation, the message you send to whichever teams you defer, and name the one deliverable you will refuse to produce in the time available.
Approach
- Sort by who owns the date. The statutory deadline is external and non-negotiable; the other two have owners who can trade scope or timing, which makes them negotiable even though both feel urgent.
- Cost each request in hours including the slow parts rather than the query: the statutory extract needs a validation pass and review time, the backlog projection needs arrival rates and a real capacity ceiling from dim_unit, and the equity ratio needs request-type stratification and a boundary-vintage-correct join to dim_geography.
- Allocate with review time inside the estimate, not after it. Finish the statutory report early enough that a second person can check it, because a late correction on a statutory return costs more than every other item on the list combined.
- Scope the backlog projection down to something honest: a capacity-bounded projection that states the stability condition, since a queue only clears when the arrival rate is below the throughput ceiling, and Little's Law relates work in progress to arrival rate times time in system only in steady state, which a seasonal arrival pattern violates.
- Turn the deferral into a commitment: a date, a named smaller interim artefact, and the specific input you need from them in the meantime, so "deferred" is falsifiable rather than a soft no.
- Refuse the one thing that cannot be done well: a single-date backlog clearance forecast with no capacity assumption, and say why in one sentence rather than negotiating it down.
Follow-up
- The operations director escalates to your manager. What do you send, and what do you not say?
- On day six the statutory extract fails a validation check. What drops, and who finds out first?
- 01
How do you handle receiving intense, highly critical feedback on your work from peers or senior researchers?
- 02
A programme director stops you and asks, "how many people are we missing?" You have fact_application, dim_constituent and dim_geography and nothing else. Before writing any SQL you owe her a written scope in thirty minutes: the question restated as something estimable, the denominator you will build and where it comes from, what the data cannot tell her, and a delivery date. Produce that scope note in under 300 words, plus the three clarifying questions whose answers would change the estimate.
- 03
Three requests land on Monday and you have one analyst-week. A statutory report of decision counts by program_code is due in nine days. An operations team wants a backlog projection to size next quarter's staffing. An equity lead wants the first-response equity ratio by deprivation quartile for a briefing in six weeks. Produce the allocation, the message you send to whichever teams you defer, and name the one deliverable you will refuse to produce in the time available.
Is this an official Institute for Defense Analyses interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Institute for Defense Analyses. Rounds and questions reflect what candidates have reported, not a process Institute for Defense Analyses has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How academic is the interview process at the Institute for Defense Analyses?
The process is highly academic. You should expect your technical panel and onsite presentation to feel very similar to a thesis defense. The interviewers are looking for deep methodological rigor, intellectual honesty, and the ability to handle constructive, highly detailed criticism of your work.
PracHub interview research ↗What should I expect during the 6-hour site visit?
The site visit is an intensive day that includes a 25-minute presentation of your work sample followed by Q&A, a technical panel with researchers, a lunch with peer-level scientists, and individual meetings with senior researchers and division directors. It is designed to evaluate both your technical depth and your cultural fit within a collaborative research environment.
PracHub interview research ↗How should I handle the intense questioning during the technical panels?
Remain calm, objective, and professional. The intense questioning is not a sign that you are failing; rather, it is how the researchers evaluate your critical thinking and how you defend your scientific choices. If you made a specific assumption in your model, explain the trade-offs honestly.
PracHub interview research ↗Are letters of recommendation and transcripts mandatory?
Yes. Because the Institute for Defense Analyses operates similarly to an academic research institute, they require official transcripts and letters of recommendation to assess your academic background and research capabilities before moving you forward in the process.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22