As a Data Scientist at Optum Health Services, you are positioned at the intersection of advanced analytics and large-scale healthcare delivery. Your work directly influences the efficacy of health interventions, clinical decision-making, and the optimization of operational workflows across one of the largest health organizations in the United States. You will tackle complex, high-dimensional datasets to derive insights that improve patient outcomes and drive business strategy.
The role is both challenging and intellectually rewarding, requiring you to translate ambiguous business problems into rigorous statistical models and actionable machine learning solutions. Whether you are working on predictive modeling for patient risk, health economic analysis, or optimizing resource allocation, your contributions have a tangible, real-world impact. You will operate in a collaborative, professional environment where technical precision meets the necessity of clear, cross-functional communication.
Preparation focus
editorialNo round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.
What to demonstrate
- Breadth across SQL, experimentation and product reasoning
- Ability to state assumptions before choosing a method
How to prepare
- Drill the practice exercises below and time yourself
- Prepare three quantified stories about decisions you drove
PracHub editorial advice for the preparation topics above.
Rates built on member counts rather than exposure
Members join and leave mid-period, so dividing events by distinct members mixes a person covered for 30 days with one covered for 365. New joiners also have artificially low observed utilisation because their claims have not arrived yet and because care takes time to initiate. Denominators must be member-months or member-years, and comparative quality measures usually need a continuous-enrolment requirement with an explicit allowable gap, stated in days.
Reading the most recent months of a claims-based series as real
Claims incur before they are reported and paid, so recent incurred months are systematically undercounted until runout completes. The lag is not uniform: pharmacy adjudicates in days, professional claims in weeks, inpatient facility claims in months. That means recent data is both too low and mix-shifted toward cheap services, which reads as a cost improvement and a utilisation drop at once. The fix is to hold the last three incurred months back or apply completion factors, and to state the paid-through date on every chart.
SQL that silently fans out on a one-to-many join
State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.
Solving silently instead of narrating the reasoning
Say which branch you are taking and why you chose it over the alternative, for example checking the denominator first because it changes what the comparison means. A correct answer that arrives with no visible path scores below a rigorous one that needed a hint.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the trade-offs between different machine learning algorithms w…
Explain the trade-offs between different machine learning algorithms when building a model for patient risk stratification.
Approach
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
Build a claims lag triangle and complete recent months
medical_claim_line has claim_id, claim_version, frequency_code, service_start_date, paid_date, allowed_amount, claim_status. Using 24 fully mature incurred months, build a development triangle of allowed dollars by incurred month and payment lag in whole months, derive cumulative completion factors by chain ladder, then estimate ultimate allowed for the three most recent incurred months at a stated paid-through date. Deduplicate to surviving claim versions before building the triangle. Return incurred_month, paid_to_date, completion_factor, estimated_ultimate.
Approach
- Deduplicate to surviving claim versions and keep paid lines first. A triangle built on all versions develops on adjustment churn rather than on payment timing, and replacements arrive late, so the distortion concentrates in exactly the tail you are trying to estimate.
- Compute lag as a whole-month difference between incurred month and paid month, not as day difference divided by 30. Ragged month lengths otherwise shuffle identical claims between lag buckets depending on which month they fall in.
- Pivot to incurred_month by lag, cumulate along lag, then form age-to-age factors as the ratio of the column L+1 total to the column L total, summed over incurred months mature enough to have both. Volume weighting is the chain ladder; a simple mean of per-month ratios lets one low-volume month dominate the factor.
- Chain the age-to-age factors from each lag to ultimate and invert: completion factor at lag L is the reciprocal of the product of factors from L onward. Apply as estimated_ultimate = paid_to_date / completion_factor.
- State the assumption you have just made, which is that the development pattern is stable. It is not, after a claims-system migration, a network change or a processing backlog, and the most recent month's factor is the least reliable because it rests on the fewest observations while carrying the largest adjustment.
Follow-up
- A processing backlog means the last two months developed slower than history. What does your estimate do, and how would you detect that before reporting?
- Pharmacy, professional and inpatient facility develop on different clocks. How do you split the triangle, and what breaks if the service mix shifts?
- What happens to a PMPM series if someone reports an unadjusted recent month, and in which direction does the error point?
Four data-quality rules over corrected and cancelled lab results
lab_result has result_id, order_id, patient_id, loinc_code, value_numeric, value_text, units, reference_low, reference_high, specimen_collected_ts, resulted_ts, result_status (preliminary, final, corrected, cancelled), supersedes_result_id. Return one row per rule with a failing count, a failing share, and one example result_id, for these rules: more than one surviving row per (order_id, loinc_code) after resolving corrections and dropping cancelled; resulted_ts earlier than specimen_collected_ts; a numeric result whose unit is not the modal unit for its loinc_code; value_numeric and value_text both null on a row with result_status 'final'.
Approach
- Resolve the correction chain before anything else. Drop result_status 'cancelled', then within each (order_id, loinc_code) keep the row whose result_id appears in no other row's supersedes_result_id. That terminal node is the surviving value and the rule holds even if a correction is loaded with a backdated timestamp.
- Build each rule as a boolean mask over a named frame rather than as a filtered copy, so the failing share has an explicit, per-rule denominator. Rule three is eligible only on rows with a non-null value_numeric; rule four only on final rows.
- For the unit rule, derive the modal unit per loinc_code from the data instead of hardcoding an expected unit. The failure being hunted is one analyte reported in two units whose values differ by a fixed factor, which no range check catches because both sets of numbers look plausible.
- Emit a tidy frame with columns rule, eligible_rows, failing_rows, failing_share, example_result_id. A printed report cannot be diffed between runs; a frame can be stored and alerted on.
- Order the output by failing_share descending so the run has a lead finding rather than four equal-weight lines.
Worked solution 25 min
- Drop cancelled rows, then compute the set of result_ids referenced by supersedes_result_id and keep rows not in that set.
- Rule one: group the survivors by (order_id, loinc_code) and count rows above 1.
- Rule two: mask on resulted_ts < specimen_collected_ts over rows where both timestamps are present.
- Rule three: compute the mode of units per loinc_code over rows with non-null value_numeric, then mask rows whose unit differs from it.
- Rule four: mask rows with result_status 'final' and both value columns null. Assemble the four results into one frame with counts, shares and an example id.
Follow-up
- The modal-unit rule fires on 8 percent of one analyte. How do you decide whether that is a unit-conversion bug or a second legitimate assay?
- What would you add to catch a value that is inside its reference range but physiologically impossible?
- Which of these four rules should block a downstream pipeline and which should only alert, and why?
Reconstruct drug coverage intervals and compute proportion of days covered
pharmacy_claim holds fill_id, member_id, ndc_code, therapeutic_class_code, fill_date, days_supply, reversal_flag and reversed_fill_id. Early refills overlap, so days_supply must be laid end to end: each fill's coverage starts at the later of its fill_date and the previous interval's end. Both rows of a reversed pair are excluded. For one therapeutic_class_code over 2025-01-01 to 2025-06-30, return per member the proportion of days covered, defined as covered days divided by days from first fill_date to window end, plus the share of members at or above 0.80. Require at least two fills and 91 days of follow-up.
Approach
- Clean the fills first. Drop rows with reversal_flag TRUE and also the fills they point at through reversed_fill_id. Dropping only the reversal row leaves an original fill counted as exposure that never reached the member. Use NOT EXISTS for that second step, since reversed_fill_id is nullable and NOT IN would return nothing.
- Shift, never stack. The shifted end is covered_end_i = GREATEST(fill_date_i, covered_end_{i-1}) + days_supply_i, which looks recursive but has a closed form: with cum_i the running sum of days_supply through fill i, covered_end_i = cum_i + MAX over j <= i of (fill_date_j - cum_{j-1}). That is SUM(days_supply) OVER (ORDER BY fill_date ROWS UNBOUNDED PRECEDING) plus a running MAX of an anchor column, so it is two window functions and no recursive CTE.
- Derive covered_start_i as GREATEST(fill_date_i, LAG(covered_end)). The intervals are disjoint by construction, so covered days is a plain sum of lengths and never needs a distinct day grid.
- Clip intervals to the window end before summing, and decide explicitly whether to clip at disenrollment as well. Supply that runs past the window must not inflate the numerator.
- Use the denominator the spec names, first fill_date to window end, and hold it constant across classes. A fixed-window denominator yields a different and usually lower number, and mixing the two across classes makes the comparison meaningless.
- Apply the inclusion rules at member level after the intervals are built, then compute the share at or above 0.80 with the qualifying member count beside it.
Follow-up
- A member is hospitalised for twelve days mid-window, when the facility supplies medication. Should those days count as covered, and what do published adherence measures do about them?
- Someone wants to classify adherence over twelve months and then compare mortality from month zero. What is wrong with that design and what fixes it?
- Two different ndc_codes in the same therapeutic class overlap. Is that stacking or switching, and how does your interval logic treat each case?
Turn overlapping coverage spans into fractional member-month denominators
member_enrollment holds span_id, member_id, plan_id, product_type, effective_date and termination_date, which is NULL while the span is active. A member can hold several spans in a year, and back-dated plan changes make spans overlap. Return one row per calendar month of 2025 with total member-months, where a member contributes covered days in the month divided by that month's length, and a day covered by two overlapping spans counts once. Both span endpoints are inclusive. Collapse overlaps with window functions before joining a month calendar, and do not assume one span per member.
Approach
- Normalise the open end first. COALESCE(termination_date, DATE '9999-12-31') or clip to the reporting end, because a NULL termination_date makes every BETWEEN comparison unknown and drops exactly the members who are still covered.
- Merge overlapping spans per member before the calendar join. Order by effective_date, carry a running MAX of prior termination_date with a frame ending one row before, and open a new island where effective_date exceeds that running max plus one day, which keeps adjacent spans contiguous rather than splitting them.
- Clip each merged island to the reporting year, then cross join to a twelve-row month calendar and keep pairs that intersect.
- Compute covered days as LEAST(island_end, month_end) minus GREATEST(island_start, month_start) plus one, and divide by the actual length of that month rather than a flat 30, so February and the 31-day months are not distorted.
- Sum the fractions by month. Merging after the calendar join also works but forces a DISTINCT over member-days, which stops scaling at population size.
Worked solution 30 min
- CTE bounded: clip spans to 2025-01-01 through 2025-12-31 with COALESCE on termination_date, discarding spans that do not intersect the year.
- CTE merged: gaps-and-islands per member using LAG and a running MAX over preceding rows to collapse overlapping and adjacent spans.
- CTE months: twelve rows of month_start, month_end and days_in_month.
- Join merged to months on interval intersection and compute covered_days per pair.
- SUM(covered_days / days_in_month::numeric) grouped by month, ordered by month.
Follow-up
- Where does a continuous-enrolment filter with an allowable 45-day gap belong, and how does that change the island boundary rule?
- One human appears under two member_ids after a master person index merge. What does that do to member-months, and how would you detect it?
- How do you treat a span that terminates on the day it starts, and what does that row usually represent?
How do you ensure your statistical analysis accounts for potential bia…
How do you ensure your statistical analysis accounts for potential biases in healthcare data?
Approach
- Clarify what is being asked and what a complete answer would contain.
- Work from the decision backwards to the evidence you would need.
- Say what you would check first and why it is the highest-information step.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
What experience do you have with peer-reviewed research or applying st…
What experience do you have with peer-reviewed research or applying statistical rigor to medical data?
Approach
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
How would you handle missing data in a clinical dataset where the abse…
How would you handle missing data in a clinical dataset where the absence of a value might be informative?
Approach
- Work from the decision backwards to the evidence you would need.
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Estimate a programme effect at a fixed risk-score enrolment cutoff
Care management enrols members whose prospective_risk_score in member_enrollment crosses a fixed annual threshold of 1.75; above the line enrolment is near-automatic and below it unavailable, and about 85 percent of members above the line actually enrol. Leadership refused a randomised holdout and wants the programme's effect on next-year allowed PMPM. Two years of scores and claims are available. State the design, the estimand it identifies, the assumptions, and the diagnostics you would run before reporting anything.
Approach
- Propose a fuzzy regression discontinuity on prospective_risk_score at 1.75. Uptake jumps but is not complete, so the estimator is the ratio of the jump in outcome to the jump in enrolment probability at the cutoff, which is a Wald ratio.
- State the estimand honestly up front: a local effect for members near 1.75, who are the least sick enrollees. It does not identify the effect on the top risk decile, which is usually the number leadership actually wants, and saying so before the result lands is what keeps the analysis credible.
- Say explicitly why the obvious alternative fails. A pre-post on enrolled members selects on a high score in year one, and such cohorts regress toward the mean in year two whether or not anything is done, so a pre-post design reports savings every time. The discontinuity is immune because it compares members on either side of an arbitrary line.
- Specify the estimation: local linear fits on each side, triangular kernel, a data-driven MSE-optimal bandwidth, and robust bias-corrected inference rather than a naive local-linear confidence interval. Do not fit a high-order global polynomial.
- Run the falsification suite before looking at the effect: a density test for manipulation of the running variable, continuity of pre-determined covariates at the cutoff, sensitivity across bandwidths, and placebo cutoffs away from 1.75.
Worked solution 30 min
- Plot the outcome and the enrolment rate against the centred running variable in uniform bins to confirm a visible jump in both.
- Estimate the first stage (jump in enrolment probability at 1.75) and the reduced form (jump in next-year PMPM), then take the ratio for the fuzzy RD estimate.
- Select the bandwidth by an MSE-optimal rule with a triangular kernel and local linear fits, and report robust bias-corrected intervals.
- Run the density test at 1.75 and the covariate continuity tests on age, prior-year cost, dual_eligible_flag, and product_type.
- Re-estimate at 0.5x, 1x, and 2x the selected bandwidth, and at placebo cutoffs of 1.50 and 2.00.
Follow-up
- Risk scores are built from coded diagnoses, and coding intensity is something people can influence. What would manipulation look like in the density, and what do you do if you see it?
- The MSE-optimal bandwidth leaves 900 members. How do you decide whether to report an underpowered estimate at all?
- Design an encouragement version instead: randomise the invitation rather than enrolment. What does instrumental variables buy you, and which assumption is most at risk?
First-pass denial rate doubled in one received week
First-pass denial rate jumped from 6.1 to 11.8 percent in a single received_date week and stayed at the new level. Nothing in the adjudication rules changed that week. You have medical_claim_line: claim_id, claim_version, frequency_code, claim_status, denial_reason_code, member_id, received_date, adjudicated_at, billing_provider_npi. The metric counts lines whose first adjudication returns 'denied', over lines received in the window at claim_version = 1. Find the cause, say whether denial behaviour actually changed, and deliver the corrected series plus the query change that fixes it.
Approach
- Cut the excess by denial_reason_code before anything else. A behaviour change spreads across codes; an upstream failure concentrates in one. If a single eligibility-related code carries nearly all of the 5.7 point excess, the claims are fine and the eligibility data was not.
- Join the denied lines to member_enrollment and compare the span's created_at to the claim's received_date. Lines denied for members whose coverage span loaded after the claim arrived are a load-ordering failure, not a coding failure.
- Check resolution downstream: take those claim_ids and look for a later version with frequency_code 7 that adjudicates to 'paid'. A high resubmit-and-pay share confirms the denial was transient and recoverable.
- Verify the denominator is not the mover. Count lines at claim_version = 1 by received_date and check the weekday mix, because submission batching makes received_date strongly day-of-week patterned and a short or holiday week shifts the denominator without changing behaviour.
- Confirm runout: the metric should be held until 60 days of adjudication have passed, so check that the comparison weeks are equally mature rather than comparing a settled week to a fresh one.
- Restate the series with the affected reason code broken out as its own line, and change the query to carry denial_reason_code into the reported breakdown rather than only the headline rate.
Follow-up
- The affected lines all get paid on resubmission. Is the first-pass denial rate then wrong, or right and uninteresting?
- How would you monitor for this class of failure without waiting for someone to notice a chart?
- What changes if the excess is spread evenly across denial_reason_code instead of concentrated?
Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and set a baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
- Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
- Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.
Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Worked solution ↗02One query pattern, written three times
- Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
- On the third attempt, write the grain of every CTE as a comment before writing its body.
- Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.
Deliverable: Three independent versions of the same query plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗03Only the statistics you will be asked to defend
- Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
- Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
- Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.
Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.
Practice prompt ↗Practice prompt ↗04One case, and the assumptions holding it up
- Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
- Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
- Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.
Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
- Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in a longer weekend block
- Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
- Re-read only your own notes from this week, and open no new material.
- Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Nearly every data role forces a trade between the analysis you want and the one that fits the decision window. Prepare a case where you deliberately shipped something less rigorous, named the weakness to the person relying on it, and said what would change your answer. The naming is the part interviewers listen for.
Describe a time you had to explain a complex technical finding to a no…
Describe a time you had to explain a complex technical finding to a non-technical stakeholder in a clinical setting.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Explain an incomplete cost series to a non-technical executive
An executive opens a dashboard showing risk-adjusted allowed PMPM down 6 percent across the last three incurred months and asks whether the cost initiative is working. You know pharmacy adjudicates within days, professional claims within weeks, and inpatient facility claims over months, so those three months are both undercounted and mix-shifted toward cheap services. There is also a genuine movement of roughly 1.5 percent in the older, complete months. Deliverable: a two-minute spoken answer, the change you make to the chart, and what you commit to telling them and when.
Approach
- The probe is whether you can be honest about uncertainty without sounding evasive. Executives read hedging as not knowing, so the answer has to end in a commitment.
- Separate the two claims immediately. The 6 percent is an artefact of incomplete data, and there is a smaller real movement in the months that are complete. Delivering only the debunk leaves them holding nothing.
- Explain the lag in their terms: the bills for the most expensive care arrive last, so an incomplete month always looks cheap and always looks like it is improving. Keep the phrase completion factor out of the first pass.
- Fix the chart instead of explaining around it. Shade or omit the incomplete months, print the paid-through date on the axis, and overlay the same series as it looked at equivalent maturity a year earlier so the shape is comparable rather than merely lower.
- Close with a date and a magnitude. Commit to a restated figure once runout matures, say roughly how far you expect the 6 percent to shrink, and name what you would need to see to call the initiative working.
Follow-up
- They ask for your best guess today, knowing it is provisional. What do you say?
- The restated number comes back at 1 percent. How do you handle having flagged the 6 percent at all?
- How do you stop this dashboard producing the same conversation next quarter?
Defend a null result against a programme sponsor
A care management programme enrolled the top 1 percent of members by prior-year allowed spend. The sponsor's deck shows allowed PMPM for enrollees falling 34 percent from the year before enrollment to the year after, and asks for budget to triple the programme. You rebuild the evaluation with a concurrent comparison group selected by the identical spend rule in the same period. The difference-in-differences estimate is a 3 percent reduction with a confidence interval spanning zero. Deliverable: how you present this, to whom, in what order, and what you propose next.
Approach
- Name the probe: whether you can deliver a finding that costs somebody their programme without softening it into uselessness or creating an adversary who routes around you next time.
- Lead with the mechanism, not the verdict. Show the comparison group's own unadjusted drop, which will be large, because a cohort selected on an extreme of the outcome regresses toward the mean whether or not anyone intervenes. The sponsor's 34 percent is mostly that, and it is a property of the selection rule rather than a criticism of their clinicians.
- Give the sponsor the finding privately before it appears in any deck their leadership sees. Being surprised in a room is what turns a methods disagreement into a political one.
- State the estimate with its interval and say what it rules out as well as what it fails to establish. A 3 percent point estimate whose interval crosses zero is not evidence of no effect; it is insufficient power to separate a modest effect from none, and those are different claims.
- Arrive with a design rather than only an objection. Propose a regression discontinuity at the enrollment threshold if the rule is applied sharply, or a randomised rollout across the next wave of eligible members, and state the sample size needed to detect the effect size the sponsor believes in.
Follow-up
- The sponsor says withholding the programme from a comparison group is unethical. What do you propose instead?
- Leadership wants one number for the board next week. What do you give them?
- What result would change your mind and make you believe the 34 percent?
- 01
Describe a time you had to explain a complex technical finding to a non-technical stakeholder in a clinical setting.
- 02
An executive opens a dashboard showing risk-adjusted allowed PMPM down 6 percent across the last three incurred months and asks whether the cost initiative is working. You know pharmacy adjudicates within days, professional claims within weeks, and inpatient facility claims over months, so those three months are both undercounted and mix-shifted toward cheap services. There is also a genuine movement of roughly 1.5 percent in the older, complete months. Deliverable: a two-minute spoken answer, the change you make to the chart, and what you commit to telling them and when.
- 03
A care management programme enrolled the top 1 percent of members by prior-year allowed spend. The sponsor's deck shows allowed PMPM for enrollees falling 34 percent from the year before enrollment to the year after, and asks for budget to triple the programme. You rebuild the evaluation with a concurrent comparison group selected by the identical spend rule in the same period. The difference-in-differences estimate is a 3 percent reduction with a confidence interval spanning zero. Deliverable: how you present this, to whom, in what order, and what you propose next.
Is this an official Optum Health Services interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Optum Health Services. Rounds and questions reflect what candidates have reported, not a process Optum Health Services has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews?
The technical interviews are generally considered average in difficulty, focusing on practical application rather than theoretical trivia. Expect to discuss your past projects in depth and solve problems that reflect real-world data challenges.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates distinguish themselves by showing both technical depth and a clear understanding of the healthcare business context. They can articulate not just the "how" of their models, but the "why" in terms of business and patient value.
PracHub interview research ↗What is the team culture like?
The culture is professional, structured, and focused on collaboration. You will be working with smart, reasonable colleagues, but you should be prepared for a environment where diverse perspectives and clear communication are highly valued.
PracHub interview research ↗How long does the process take?
While timelines vary by team, the process is generally structured and moves at a steady, professional pace. Ensure you are prepared for each stage by reviewing your past projects and reflecting on how your skills align with the requirements of the role.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22