Washington University in St. Louis · Data Scientist
Updated · 2026-09-22

Washington University in St. Louis Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Washington University in St. Louis, you are positioned at the cutting edge of clinical research and healthcare innovation. This role is deeply embedded in the university's world-renowned medical campus, specifically focusing on clinically-driven medical analytics and spine health. You are not just crunching numbers; you are directly influencing patient care, surgical outcomes, and the future of orthopedic and neurological treatments.

In modelling rounds the live question is usually why this model class for this problem, and how you would know six weeks after launch that it is still working. Deriving gradients by hand is rarely what is being probed.

Washington University in St. Louis candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Reconstruct active time from raw heartbeat streamsModel seat activation as the renewal leading indicatorAlign cohorts to term weeks, not calendar weeks

29 min read

Practice 13 Data Scientist prompts
13Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Washington University in St. Louis, you are positioned at the cutting edge of clinical research and healthcare innovation. This role is deeply embedded in the university's world-renowned medical campus, specifically focusing on clinically-driven medical analytics and spine health. You are not just crunching numbers; you are directly influencing patient care, surgical outcomes, and the future of orthopedic and neurological treatments.

Your work bridges the gap between raw physiological data and actionable clinical insights. By analyzing complex streams of data from wearables and electronic health records (EHR), you will help clinicians understand patient recovery trajectories, predict surgical complications, and optimize spine health interventions. The impact of this position is profound, as your models and analyses will directly inform the decisions made by top-tier surgeons and medical researchers.

Expect a highly collaborative, academically rigorous, and mission-driven environment. values data scientists who can handle the immense scale and complexity of medical data while maintaining a strong focus on patient outcomes. You will tackle unique challenges, such as dealing with noisy time-series data from consumer wearables, ensuring strict privacy compliance, and translating complex machine learning concepts to non-technical medical professionals.

01

Recruiter Screening

reported

Most candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.

What to demonstrate

  • Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
  • Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
  • The substance of the questions you ask back, which an experienced screener reads as a level signal

How to prepare

  • Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
  • Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
  • Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
PracHub interview research ↗
02

Technical Screening

reported

Much of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.

What to demonstrate

  • Whether the query you write matches the plan you just described
  • What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
  • Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly

How to prepare

  • Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
  • Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
  • Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
PracHub interview research ↗
03

Panel Interview

reported

A day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.

What to demonstrate

  • Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
  • Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
  • Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method

How to prepare

  • Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
  • Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
  • Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
  • Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Learner-level standard errors on class-level interventions.

Anything an instructor controls, and anything deployed by school, is assigned at the section or school level, and outcomes within a section are correlated through the shared instructor, schedule, and device fleet. Treating the learner as the unit of independence understates variance by the design effect 1 + (m - 1) * rho for roughly equal cluster sizes. At a typical section size of 25 and rho of 0.15 that is a factor near 4.6 on variance, which routinely converts a null into a 'significant' result.

02

The academic calendar creates structural breaks that look like product effects.

Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.

03

Reading experiment results before checking the arm split

Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.

04

Reading a dozen metrics with no multiplicity control

Nominate one primary metric before launch and treat the rest as guardrails or exploratory, with Bonferroni or Benjamini-Hochberg applied when you intend to make claims from them. Twenty independent tests at 0.05 under the null produce at least one false positive about 64 percent of the time.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

10 technical prompts3 include a worked solution

What are the risks of using deep learning for clinical decision suppor…

medium
statistics and probability

What are the risks of using deep learning for clinical decision support, and how would you mitigate them?

Approach
  1. Translate the result into the decision it informs, in one plain sentence.
  2. Write down the assumption the method needs before you use the method.
  3. Sanity-check the answer against a simple bound or a simulated case.
Follow-up
  • How would you explain this result to someone who does not know statistics?
  • Which assumption here is most likely to be violated in practice?

How do you evaluate the performance of a classification model when you…

medium
machine learning and modelling

How do you evaluate the performance of a classification model when your positive class (e.g., surgical complications) is extremely rare?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Say how the offline result would be validated online before it is trusted.
  3. Check what information would not exist at prediction time, and exclude it.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

How would you build a model to predict the likelihood of a patient req…

medium
machine learning and modelling

How would you build a model to predict the likelihood of a patient requiring revision spine surgery within one year?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Say how the offline result would be validated online before it is trusted.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • Where could label leakage enter this setup?
  • What would you monitor after launch to know the model is still valid?

Explain SHAP values to me as if I were a surgeon with no machine learn…

medium
machine learning and modelling

Explain SHAP values to me as if I were a surgeon with no machine learning background.

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Pick an evaluation metric that matches the cost of each error type, not a default.
  3. Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • How would you choose the decision threshold, and who owns that choice?

Sessionise a raw heartbeat stream into activity rows

hardWorked solution
sessionisationevent streamsgroupbytime series

You get a DataFrame heartbeats with columns learner_id, content_item_id, content_version_no, and ts (UTC timestamps, unsorted). Produce one row per learner per session per content item with started_at, ended_at, wall_seconds, and active_seconds. A session closes after 30 minutes with no heartbeat from that learner. active_seconds sums consecutive gaps inside a session, capping each gap at 120 seconds. Attribute each gap to the content item of the earlier heartbeat. State your rule for a gap of exactly 1800 seconds and for a session with a single heartbeat.

Approach
  1. Sort by learner_id then ts, and compute the gap with a per-learner diff so the first heartbeat of each learner has no gap. A global diff lets one learner's last heartbeat set the next learner's first gap.
  2. Flag a session break where gap > 1800 seconds, then cumsum that boolean within the learner to get a session index. Say whether exactly 1800 opens a new session; either choice is defensible but it must be stated because it moves the session count.
  3. Each gap covers the interval before the current heartbeat, so attribute min(gap, 120) to the content_item_id of the previous row via shift, not the current row.
  4. Group by (learner_id, session_index, content_item_id): sum the attributed capped gaps for active_seconds, take min and max ts for started_at and ended_at, and derive wall_seconds from those two.
  5. Handle the degenerate group explicitly. A single heartbeat gives active_seconds 0 and wall_seconds 0. Dropping those rows quietly removes real opens and biases any completion or abandonment rate computed later.
  6. Validate that within every session the summed active_seconds never exceeds the session's wall_seconds, which catches attribution and cap errors in one assertion.
Worked solution 35 min
  1. Sort by (learner_id, ts) and compute gap = ts.diff().dt.total_seconds(), then null the gap on the first row of each learner using a learner-change mask.
  2. new_session = gap.isna() | (gap > 1800); session_index = new_session.groupby(learner_id).cumsum().
  3. contribution = gap.clip(upper=120).fillna(0), attributed to the previous row's content_item_id using shift within the learner and session.
  4. Aggregate by (learner_id, session_index, content_item_id) with sum for active_seconds and min and max of ts for the timestamps.
  5. Assert per session that sum(active_seconds) <= (max(ts) - min(ts)) in seconds, and print the count of single-heartbeat groups.
EXPECTED RESULTA frame keyed by (learner_id, session_index, content_item_id) where active_seconds equals the sum of min(gap, 120) over consecutive heartbeats with the first heartbeat contributing zero, wall_seconds equals ended_at minus started_at, and sum(active_seconds) <= wall_seconds holds for every session.
Follow-up
  • The client emits a heartbeat every 60 seconds on desktop and every 15 seconds on a managed device fleet. What does the 120-second cap do to the comparison between those two populations?
  • Your reconstructed active_seconds disagrees with the stored column by 4 percent overall. How do you localise the disagreement rather than argue about the total?
  • A learner has two tabs open on different items at once. What does your output claim, and what should it claim?

For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Design one test end to end on paper
  • Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
  • Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
  • State in advance what you will do if the primary metric is flat while a secondary metric is significant.

Deliverable: A one-page test design with a decision rule written before launch.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Power arithmetic until it is automatic
  • Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
  • Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
  • Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.

Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.

Practice prompt ↗Practice prompt ↗
03Variance and the unit-of-analysis problem
  • Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
  • Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
  • Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.

Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.

Practice prompt ↗Practice prompt ↗
04Validity threats you can actually test for
  • Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
  • Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
  • Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.

Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05When randomization is not available
  • Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
  • Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
  • List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.

Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.

Practice prompt ↗Practice prompt ↗
06The readout query
  • Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
  • Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
  • Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.

Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.

Practice prompt ↗Practice prompt ↗
07Present it to someone who will not read the appendix
  • Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
  • Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
  • Rewrite your opening line so the recommendation lands before any methodology.

Deliverable: A one-page readout whose first line is the recommendation.

Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.

How do you ensure your work remains focused on patient outcomes rather…

medium
behavioural and stakeholder questions

How do you ensure your work remains focused on patient outcomes rather than just technical novelty?

Approach
  1. Close with what you would do differently, concretely.
  2. Quantify the outcome, including what you would not claim credit for.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

How do you handle missing or irregular time-series data from a patient…

medium
behavioural and stakeholder questions

How do you handle missing or irregular time-series data from a patient's wearable device?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Name the disagreement or constraint, and how you resolved it with evidence.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • How did you know the outcome was caused by your change?
  • What would you do differently if you ran that project again?

Defend a null result against an already-announced launch

medium
null resultsexperiment designstakeholder communication

A section-randomised test of a new practice-sequencing feature ran across 62 class sections in 11 orgs. Primary metric: verified mastery events per active learner. With section-clustered standard errors the effect is +1.8% (95% CI -4.1% to +7.9%). Leadership has already told two district administrators that the feature improves outcomes. You have ten minutes in a review. Deliverable: what you say, what you put on one slide, and what you propose next. Probed: whether you can hold a null under commercial pressure without either caving or lecturing the room on statistics.

Approach
  1. Lead with the decision, not the p-value: state the interval in the unit the audience already uses (mastery events per active learner per week) and say plainly what the study could and could not have detected.
  2. Separate the interval's half-width from the minimum detectable effect before you put either on a slide. The stated 95% CI of -4.1 to +7.9 has a half-width of 6.0 points, so the clustered standard error is 6.0 / 1.96 = 3.06 points; the MDE at 80% power and two-sided alpha = 0.05 is (1.96 + 0.8416) * 3.06 = about 8.6 points, roughly 1.43 times the half-width. Quote 8.6, and show that it follows from 62 sections at the section size and intraclass correlation you assumed via the design effect 1 + (m - 1) * rho, so that 'no effect found' is visibly separated from 'too small to see'.
  3. Check the pre-registered guardrails before conceding anything: median active_seconds per mastery event and the calibrated irt_b distribution of the items backing mastery decisions. If both are flat, the null is about learning, not about a broken pipeline.
  4. Handle the promise made to the two districts as its own item with a named owner, rather than leaving the correction implicit in a statistics discussion.
  5. Close with the cheapest decision-grade next step: extend to enough additional already-instrumented sections to reach the effect size leadership actually cares about, with the stopping rule fixed before the extension starts.
Follow-up
  • If the interval had been -0.5% to +8.0%, would you support launch, and what exactly changes in your reasoning?
  • Leadership asks you to re-run with learner-level standard errors since it is the same data. What do you say?
  • What would you have set up at design time so that this meeting was easier?
  • 01

    How do you ensure your work remains focused on patient outcomes rather than just technical novelty?

  • 02

    How do you handle missing or irregular time-series data from a patient's wearable device?

  • 03

    A section-randomised test of a new practice-sequencing feature ran across 62 class sections in 11 orgs. Primary metric: verified mastery events per active learner. With section-clustered standard errors the effect is +1.8% (95% CI -4.1% to +7.9%). Leadership has already told two district administrators that the feature improves outcomes. You have ten minutes in a review. Deliverable: what you say, what you put on one slide, and what you propose next. Probed: whether you can hold a null under commercial pressure without either caving or lecturing the room on statistics.

PracHub interview preparation framework ↗
Is this an official Washington University in St. Louis interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Washington University in St. Louis. Rounds and questions reflect what candidates have reported, not a process Washington University in St. Louis has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
Do I need a background in medicine or spine surgery to be hired?

While a medical background is not strictly required, a strong interest in healthcare and the ability to rapidly learn clinical terminology is essential. You must demonstrate that you can bridge the gap between data science and clinical application.

PracHub interview research ↗
What is the typical timeline for the interview process?

Academic and medical institution hiring can sometimes move slower than the tech sector. Expect the process from initial screen to final offer to take anywhere from 4 to 8 weeks, as coordinating schedules with busy clinical faculty can take time.

PracHub interview research ↗
Will I be expected to publish research?

Yes, contributing to academic publications is often a key component of data science roles within medical research groups at Washington University in St. Louis. Highlighting any past experience with academic writing or research presentation is highly beneficial.

PracHub interview research ↗
How much coding vs. clinical strategy will I be doing?

This is a highly technical role. You will spend the majority of your time coding (Python/R/SQL), building data pipelines, and training models. However, the strategic clinical aspect is what guides your coding, so you must be comfortable with both.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.