As a Data Scientist at the U.S. Food and Drug Administration, your work directly supports public health, safety, and regulatory science at a national scale. You will leverage complex datasets, advanced statistical methods, and machine learning models to help evaluate medical products, monitor food safety, and drive evidence-based policy decisions. Your analyses inform critical regulatory reviews and shape how millions of citizens experience health and safety safeguards.
The role sits at the intersection of rigorous statistical methodology, large-scale data engineering, and high-impact public service. You will collaborate closely with multidisciplinary teams of statisticians, epidemiologists, clinicians, and regulatory project managers. Whether you are designing rigorous experiments, auditing predictive models, or diagnosing metric anomalies in safety surveillance systems, your insights carry significant weight and operational responsibility.
Expect a mission-driven environment where scientific rigor, reproducibility, and clarity of communication are prized above all else. You will not only build complex analytical pipelines but also translate those findings into actionable recommendations for non-technical stakeholders and agency leadership. Success in this position requires a balance of deep technical mastery and the ability to communicate complex findings with absolute precision.
HR Screening
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Screening Calls
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Formal Research Presentation
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
Interactive Panel Conversations
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
Final Panel Evaluations
reportedWhere a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.
What to demonstrate
- Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
- Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
- Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
- Whether you ask what decision is actually on the table before explaining anything
How to prepare
- Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
- For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
- Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub editorial advice for the preparation topics above.
Rates built on member counts rather than exposure
Members join and leave mid-period, so dividing events by distinct members mixes a person covered for 30 days with one covered for 365. New joiners also have artificially low observed utilisation because their claims have not arrived yet and because care takes time to initiate. Denominators must be member-months or member-years, and comparative quality measures usually need a continuous-enrolment requirement with an explicit allowable gap, stated in days.
Treating clinical measurements as missing at random
A lab result, a vital sign, or a screening exists because someone ordered it, and ordering tracks suspicion of disease, visit frequency, and site workflow. Imputing the mean or dropping incomplete rows biases the population estimate and can flip the sign of an association, because the untested are systematically healthier or systematically disengaged. The presence indicator is often more predictive than the value, which is a warning sign rather than a feature win: a model that learns test ordering will not transfer to a site with different protocols.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Treating a non-significant result as proof of no effect
Say whether the confidence interval excludes the effect sizes you would have cared about. If it does not, the honest reading is that the test was underpowered, so report the minimum detectable effect the design could have found and what sample size would resolve it.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Compare parametric and non-parametric testing methods when dealing wit…
Compare parametric and non-parametric testing methods when dealing with highly skewed biomedical datasets.
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Say what the estimate is of, and over what population it generalises.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
How do you test for normality and homoscedasticity in regression resid…
How do you test for normality and homoscedasticity in regression residuals, and what remedial steps do you take when assumptions fail?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Sessionise transfer chains into episodes, then count readmissions
encounter has encounter_id, patient_id, facility_id, encounter_type, admit_ts, discharge_ts, admission_type, discharge_disposition, is_planned_admission, principal_diagnosis_code. Chain acute-to-acute transfers into episodes of care: an inpatient encounter discharged as transfer_acute, followed by another inpatient encounter for the same patient at a different facility admitting within 24 hours of that discharge, belongs to the same episode. Then compute the 30-day unplanned readmission rate over episodes, excluding index episodes that are planned, that end in expired, hospice or against medical advice, or that are still open. Return encounter_id to episode_id, plus the rate.
Approach
- Restrict to inpatient encounters for the chaining step and sort by patient_id, admit_ts. Observation and emergency encounters are not acute inpatient stays and pooling them changes both the chain and the denominator.
- Build a boolean continues flag: the previous row is the same patient, its discharge_disposition is transfer_acute, its facility_id differs from this row's, and admit_ts minus previous discharge_ts is between 0 and 24 hours. Episode id = cumsum of the negation of that flag. This is the vectorised form of a sessionisation loop and is why the pattern generalises to any event stream.
- Collapse to episode grain taking first admit_ts, last discharge_ts, last discharge_disposition, and any() of is_planned_admission. The episode inherits its exit state, not its entry state, which is why a transfer chain ending at home is one home discharge rather than two transfers.
- Apply index exclusions at episode grain. Dropping episodes that end in death is not optional: a dead patient cannot be readmitted, so leaving them in inflates the denominator and depresses the rate for exactly the panels caring for the sickest members.
- For each surviving index episode, find the next unplanned acute inpatient episode for that patient admitting within 30 days of the index discharge_ts. Count events at episode grain, since one patient can contribute several index episodes. Report the raw rate and say plainly that it is not comparable across panels without risk adjustment, which is what the observed-over-expected form exists for.
Worked solution 45 min
- Filter to encounter_type 'inpatient', sort by patient_id and admit_ts, and shift discharge_ts, discharge_disposition, facility_id and patient_id by one row.
- Compute continues as same patient AND prior disposition is transfer_acute AND facility differs AND 0 <= (admit_ts - prior discharge_ts) <= 24 hours; episode_id = (~continues).cumsum().
- Aggregate to episodes: first admit_ts, last discharge_ts, last discharge_disposition, any planned flag, patient_id.
- Drop index-ineligible episodes: planned, disposition in expired, hospice or ama, and null discharge_ts.
- For each eligible index episode, search the same patient's later episodes for an unplanned admit_ts within 30 days of index discharge_ts; a merge_asof on patient with a forward direction and a 30-day tolerance does this without a nested loop.
- Rate = index episodes with an event divided by eligible index episodes; return the encounter-to-episode map alongside it.
Follow-up
- discharge_ts is null on an open stay in the middle of a chain. What does your flag do, and what should it do?
- A patient is transferred out and then back to the originating facility within 24 hours. Should that chain, and does your facility_id condition handle it?
- Is a readmission itself eligible to serve as a later index episode? Say what you chose and what the choice does to the rate.
Given a table of user activity logs, how would you write a query to id…
Given a table of user activity logs, how would you write a query to identify returning users using conditional aggregation and self-joins?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Explain how you would use ranking and partitioning window functions to…
Explain how you would use ranking and partitioning window functions to find the top three most frequent error codes per region.
Approach
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Reconstruct drug coverage intervals and compute proportion of days covered
pharmacy_claim holds fill_id, member_id, ndc_code, therapeutic_class_code, fill_date, days_supply, reversal_flag and reversed_fill_id. Early refills overlap, so days_supply must be laid end to end: each fill's coverage starts at the later of its fill_date and the previous interval's end. Both rows of a reversed pair are excluded. For one therapeutic_class_code over 2025-01-01 to 2025-06-30, return per member the proportion of days covered, defined as covered days divided by days from first fill_date to window end, plus the share of members at or above 0.80. Require at least two fills and 91 days of follow-up.
Approach
- Clean the fills first. Drop rows with reversal_flag TRUE and also the fills they point at through reversed_fill_id. Dropping only the reversal row leaves an original fill counted as exposure that never reached the member. Use NOT EXISTS for that second step, since reversed_fill_id is nullable and NOT IN would return nothing.
- Shift, never stack. The shifted end is covered_end_i = GREATEST(fill_date_i, covered_end_{i-1}) + days_supply_i, which looks recursive but has a closed form: with cum_i the running sum of days_supply through fill i, covered_end_i = cum_i + MAX over j <= i of (fill_date_j - cum_{j-1}). That is SUM(days_supply) OVER (ORDER BY fill_date ROWS UNBOUNDED PRECEDING) plus a running MAX of an anchor column, so it is two window functions and no recursive CTE.
- Derive covered_start_i as GREATEST(fill_date_i, LAG(covered_end)). The intervals are disjoint by construction, so covered days is a plain sum of lengths and never needs a distinct day grid.
- Clip intervals to the window end before summing, and decide explicitly whether to clip at disenrollment as well. Supply that runs past the window must not inflate the numerator.
- Use the denominator the spec names, first fill_date to window end, and hold it constant across classes. A fixed-window denominator yields a different and usually lower number, and mixing the two across classes makes the comparison meaningless.
- Apply the inclusion rules at member level after the intervals are built, then compute the share at or above 0.80 with the qualifying member count beside it.
Worked solution 40 min
- Filter to the therapeutic_class_code and window, then remove reversal rows and their originals with NOT EXISTS.
- Per member ordered by fill_date, compute cum_supply and anchor = fill_date - (cum_supply - days_supply), then covered_end = cum_supply + MAX(anchor) OVER (ORDER BY fill_date ROWS UNBOUNDED PRECEDING).
- Compute covered_start = GREATEST(fill_date, LAG(covered_end) OVER (ORDER BY fill_date)), clip both ends to the window, and drop intervals that collapse.
- Sum interval lengths as covered_days, compute follow_up_days from first fill_date to window end, and divide.
- Apply the two-fill and 91-day rules, then aggregate the share at or above 0.80.
Follow-up
- A member is hospitalised for twelve days mid-window, when the facility supplies medication. Should those days count as covered, and what do published adherence measures do about them?
- Someone wants to classify adherence over twelve months and then compare mortality from month zero. What is wrong with that design and what fixes it?
- Two different ndc_codes in the same therapeutic class overlap. Is that stacking or switching, and how does your interval logic treat each case?
How would you design a metric framework to measure the effectiveness o…
How would you design a metric framework to measure the effectiveness of a new post-market safety surveillance system?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
If key engagement metrics for an internal reporting dashboard drop by …
If key engagement metrics for an internal reporting dashboard drop by fifteen percent week-over-week, how would you investigate and isolate the root cause?
Approach
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you balance competing product priorities when defining success …
How do you balance competing product priorities when defining success metrics for a regulatory review tracking tool?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
You notice an unexpected spike in data processing failures on a core p…
You notice an unexpected spike in data processing failures on a core pipeline; walk through your diagnostic framework for identifying the bottleneck.
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Explain how you account for novelty and primacy effects when running d…
Explain how you account for novelty and primacy effects when running digital product experiments.
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Decide the analysis before seeing data, including how long it runs and when you look.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
How do you determine sample size and power requirements when dealing w…
How do you determine sample size and power requirements when dealing with low-frequency, high-severity public health events?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
Choose the primary metric for a high-risk care management program
A care-management team enrolls members in the top 1 percent of prospective_risk_score and wants one number it is judged on each quarter. The proposed metric is allowed PMPM for enrolled members, built from medical_claim_line.allowed_amount and pharmacy_claim.allowed_amount over member-months from member_enrollment. Propose the primary metric, two guardrails, and the reporting cadence. In three sentences, explain why the proposed metric will show large savings in year one even if the program has no effect, and what design change removes that.
Approach
- Name the failure before proposing anything. Members are selected on an extreme value of a noisy quantity, and the selecting year captured chronic severity plus one-off events that do not repeat. The cohort's spend falls the following year whether or not anyone intervenes, so a pre-post comparison on this cohort reports savings every time.
- Quantify it rather than asserting it: compute the selection-year and following-year allowed PMPM for last year's top-1-percent cohort and for the whole book. The cohort's decline is usually several times the population's, and that gap is the free saving any pre-post design will bank.
- Rebuild the metric as a difference, not a level. Primary: the difference-in-differences in risk-adjusted allowed PMPM between enrolled members and a concurrent comparison group selected by the identical rule in the same period, drawn from capacity overflow, wait-list, or members just below the selection threshold.
- Handle the exposure denominator honestly. Member-months come from coverage spans, partial months count fractionally, and death and disenrollment are not neutral: a member who dies stops accruing cost and member-months together, so state whether decedents stay in the analysis and for how long.
- Pick guardrails that catch the two bad ways cost falls: ambulatory care sensitive admissions per 1,000 member-years, and emergency department visits per 1,000 member-years, both risk-adjusted. A cost drop from care that was avoided rather than delivered shows up in one of those two within a quarter or two.
- Set the cadence to incurred quarters reported at a stated paid-through date, withholding or completion-factoring the three most recent incurred months, and say that the quarter is restated until runout completes.
Worked solution 30 min
- Compute the regression-to-the-mean baseline: last year's top-1-percent cohort allowed PMPM in the selection year and the following year, alongside the same two numbers for the whole population. Write both as a percent change.
- Define the comparison group by the identical selection rule and the identical period, and list the three ways it can still differ from the enrolled group.
- Write the primary metric as a difference-in-differences on risk-adjusted allowed PMPM, stating the risk-adjustment method and the period over which the risk score is measured.
- Specify the two guardrails with full numerator, denominator and window, in member-years not member counts.
- State the cadence, the paid-through date, the restatement policy, and the pre-registered decision rule for what counts as a result worth acting on.
Follow-up
- The comparison group is wait-listed members. What is the argument that they are not comparable, and what would you show to answer it?
- If the difference-in-differences estimate is negative but the parallel-trends check on the two pre-periods fails, what do you report?
- Enrollment into the program takes 45 days from identification. How does that interval create immortal time, and what do you do about it?
Protocol deviations fell as new sites came online
Protocol deviations per 100 completed visits across a multi-site study fell from 4.6 to 2.9 over two months while randomisation accelerated. study_subject_visit carries study_id, site_id, subject_id, visit_number, visit_name, planned_visit_date, actual_visit_date, visit_window_days, visit_status, protocol_deviation_flag, deviation_severity, data_entry_ts and site_activation_date. Operations wants to credit a retraining rollout that began in that window. Establish whether deviation behaviour changed, and deliver a version of the series that is safe to review monthly.
Approach
- Treat deviation flags as late-arriving data. Build a lag triangle of data_entry_ts minus actual_visit_date, because deviations are largely identified during monitoring review weeks after the visit, so the two most recent months are structurally undercounted in exactly the way an improvement looks.
- Hold back the months that are not developed, or apply development factors fitted on fully developed months, and restate. Do this before any explanation involving training.
- Attack the denominator. It counts visits with visit_status 'completed', while out_of_window and missed visits are excluded even though out-of-window attendance is itself a deviation in most protocols, so a site that pushes visits out of window lowers the rate twice over.
- Standardise the visit mix. Newly activated sites contribute mostly screening and baseline visits, which carry fewer procedures and fewer opportunities to deviate than later treatment visits, so compute the rate within visit_number strata and reweight to a fixed mix before comparing months.
- Split major from minor deviation_severity and report them separately. A single major deviation can remove a subject from the per-protocol analysis set while ten minor ones may not, so a pooled rate can fall while the consequential series rises.
- Before attributing anything to training, model at the right level: subjects are nested in sites, so use a random intercept per site or cluster-robust standard errors. With average cluster size m and intracluster correlation rho the variance is understated by the design effect 1 + (m - 1) times rho, which at m = 60 and rho = 0.02 is about 2.2, so naive intervals are roughly 1.5 times too narrow.
Follow-up
- Design a defensible evaluation of the retraining rollout, given it was deployed site by site and cannot be withheld.
- Your reweighted series is flat but major deviations rose. What do you escalate and to whom?
- How would you decide the holdback period, and how would you communicate a provisional number without it being quoted as final?
For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Build a fixture you can check answers against
- Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
- Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
- Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.
Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Joins, filters and NULL semantics
- Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
- Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
- Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.
Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Window functions and frames
- Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
- Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
- Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.
Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The four analytical query patterns
- Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
- Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
- Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.
Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Write SQL the way you will have to write it live
- Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
- Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
- Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.
Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.
Practice prompt ↗Practice prompt ↗06One day for everything that is not SQL
- Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
- Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
- Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.
Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.
Practice prompt ↗Practice prompt ↗07Full loop rehearsal
- Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
- Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
- Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.
Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.
Explain the concept of statistical significance and how you communicat…
Explain the concept of statistical significance and how you communicate false positive versus false negative trade-offs to non-technical stakeholders.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Close with what you would do differently, concretely.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Scope a one-line request for a readmission rate
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
Approach
- Read the probe: this tests whether you convert ambiguity into a specification without either stalling or guessing silently. Both failure modes are common, and the silent guess is worse because nobody can see it.
- Rank your questions by leverage on the number rather than by curiosity. Whether the numerator is all-cause or unplanned, and whether the clock starts at discharge_ts, move the result far more than a tie-break rule on same-day returns.
- Ask about the comparison first. A number going to a committee will be compared to something, and whether that something is last quarter, another panel, or a national figure decides whether you owe them a raw rate, a risk-adjusted rate, or an observed-over-expected ratio.
- Commit to a default in writing so silence does not block you: unplanned acute inpatient readmission within 30 days of discharge_ts, index stays excluding planned admissions, acute-to-acute transfers, discharges against medical advice, and dispositions of expired or hospice, denominator of eligible index discharges, paid-through date stated.
- Write the caveat as a restriction on use rather than a hedge: name the one comparison the figure supports and the one it does not.
Follow-up
- They reply that they want it 'the way the benchmark does it'. What do you ask next?
- The committee wants it split by attending provider. What changes in the specification, and what do you refuse to show?
- How does your answer change if the meeting is tomorrow rather than Friday?
Defend a null result against a programme sponsor
A care management programme enrolled the top 1 percent of members by prior-year allowed spend. The sponsor's deck shows allowed PMPM for enrollees falling 34 percent from the year before enrollment to the year after, and asks for budget to triple the programme. You rebuild the evaluation with a concurrent comparison group selected by the identical spend rule in the same period. The difference-in-differences estimate is a 3 percent reduction with a confidence interval spanning zero. Deliverable: how you present this, to whom, in what order, and what you propose next.
Approach
- Name the probe: whether you can deliver a finding that costs somebody their programme without softening it into uselessness or creating an adversary who routes around you next time.
- Lead with the mechanism, not the verdict. Show the comparison group's own unadjusted drop, which will be large, because a cohort selected on an extreme of the outcome regresses toward the mean whether or not anyone intervenes. The sponsor's 34 percent is mostly that, and it is a property of the selection rule rather than a criticism of their clinicians.
- Give the sponsor the finding privately before it appears in any deck their leadership sees. Being surprised in a room is what turns a methods disagreement into a political one.
- State the estimate with its interval and say what it rules out as well as what it fails to establish. A 3 percent point estimate whose interval crosses zero is not evidence of no effect; it is insufficient power to separate a modest effect from none, and those are different claims.
- Arrive with a design rather than only an objection. Propose a regression discontinuity at the enrollment threshold if the rule is applied sharply, or a randomised rollout across the next wave of eligible members, and state the sample size needed to detect the effect size the sponsor believes in.
Follow-up
- The sponsor says withholding the programme from a comparison group is unethical. What do you propose instead?
- Leadership wants one number for the board next week. What do you give them?
- What result would change your mind and make you believe the 34 percent?
- 01
Explain the concept of statistical significance and how you communicate false positive versus false negative trade-offs to non-technical stakeholders.
- 02
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
- 03
A care management programme enrolled the top 1 percent of members by prior-year allowed spend. The sponsor's deck shows allowed PMPM for enrollees falling 34 percent from the year before enrollment to the year after, and asks for budget to triple the programme. You rebuild the evaluation with a concurrent comparison group selected by the identical spend rule in the same period. The difference-in-differences estimate is a 3 percent reduction with a confidence interval spanning zero. Deliverable: how you present this, to whom, in what order, and what you propose next.
Is this an official U.S. Food and Drug Administration interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at U.S. Food and Drug Administration. Rounds and questions reflect what candidates have reported, not a process U.S. Food and Drug Administration has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the entire interview process typically take?
The timeline can vary significantly depending on agency onboarding protocols and scheduling availability across multiple branches. While some stages move quickly, expect the entire process from application to final offer to span several weeks or even months.
PracHub interview research ↗What is the most important factor in passing the interview loop?
Demonstrating rigorous scientific thinking and clear communication is paramount. Interviewers want to know that you understand the mathematical foundations of your work and that you can explain complex analytical trade-offs simply and accurately.
PracHub interview research ↗Are there live coding exams during the loop?
While technical interviews will test your analytical reasoning and SQL query writing, the evaluation places equal or greater weight on your research presentation, domain knowledge, and behavioral alignment. Be prepared to discuss your past code and analytical choices in depth.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22