A Data Scientist at Garner Health plays a critical role in executing the company's core mission: transforming the healthcare economy by delivering high-quality, affordable care. By fundamentally reimagining how healthcare benefits are designed and utilized, Garner Health relies heavily on data-driven insights to steer members toward high-performing, cost-effective medical providers. The algorithms and models built by the data science team directly power the recommendation engine that guides patients through their healthcare journeys, making this role central to both user satisfaction and the company's business model.
In this position, you will work on highly complex, ambiguous data challenges, such as ranking and measuring doctor performance based on sparse or noisy historical patient outcome data. Because healthcare datasets are massive yet frequently incomplete, your work will involve building sophisticated statistical frameworks to account for small sample sizes, selection biases, and regional variations. Additionally, you will partner closely with Product Managers and Software Engineers to build and scale population health algorithms that proactively identify high-risk members and enable timely, customized care interventions.
Ultimately, being a Data Scientist at Garner Health requires a unique blend of rigorous statistical thinking, production-grade programming, and a product-oriented mindset. It is an opportunity to work on a high-stakes, real-world optimization problem where your algorithms have a direct, measurable impact on human health outcomes and financial affordability.
Recruiter Phone Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Screening
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Take-Home Case Study
reportedA take-home is graded as an argument, not as a notebook. Somebody reads the submission without you in the room, so every choice has to survive on the page: why the question was framed this way, and what was deliberately left out. The gap between a strong and a weak submission is almost never model quality. It is whether the writeup names the specific question it answers and commits to a recommendation, including what evidence would overturn it. A high-accuracy model attached to no conclusion reads as effort that stopped before the decision.
What to demonstrate
- Whether the question you answered is stated outright, and whether it is the question the prompt posed rather than an easier neighbour of it
- Whether the recommendation is specific enough to act on, with the uncertainty attached to it instead of parked in a caveats section at the end
- Whether analytical choices such as the metric definition, the population filter and the time window are justified in the prose, not merely visible in code
How to prepare
- Take a dataset you have already worked with, write the one-paragraph conclusion first, then check whether the analysis you were planning actually supports it and cut whatever does not
- Practise stating a metric in one sentence that fixes the population, the time window and the denominator, then confirm your query computes exactly that sentence and nothing adjacent to it
- Hand a draft to someone outside the problem and ask them to tell you back what you recommended and why; anything they cannot recover is not on the page yet
Panel Interviews
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
1 candidate reports. Individual accounts describe a particular role and hiring cycle.
Garner Health Software Engineer Interview Experience — Scaling an Appointment Booking System
My Garner Health interview was a system design round. The first question was to design an appointment booking system where patients could search for doctors and book appointments. I discussed the core Patient, Doctor, and Appointment data models; how to query a doctor's available time slots; and how to book or cancel an appointment. The design also needed to prevent multiple patients from booking…
Read full experiencePracHub editorial advice for the preparation topics above.
Treating clinical measurements as missing at random
A lab result, a vital sign, or a screening exists because someone ordered it, and ordering tracks suspicion of disease, visit frequency, and site workflow. Imputing the mean or dropping incomplete rows biases the population estimate and can flip the sign of an association, because the untested are systematically healthier or systematically disengaged. The presence indicator is often more predictive than the value, which is a warning sign rather than a feature win: a model that learns test ordering will not transfer to a site with different protocols.
Reading the most recent months of a claims-based series as real
Claims incur before they are reported and paid, so recent incurred months are systematically undercounted until runout completes. The lag is not uniform: pharmacy adjudicates in days, professional claims in weeks, inpatient facility claims in months. That means recent data is both too low and mix-shifted toward cheap services, which reads as a cost improvement and a utilisation drop at once. The fix is to hold the last three incurred months back or apply completion factors, and to state the paid-through date on every chart.
Explaining an aggregate move without decomposing the mix shift
Split the change in the aggregate into within-segment movement and movement in segment weights before you explain it. Every segment's rate can fall while the overall rate rises, purely because volume shifted toward segments that already had higher rates.
Reaching for a model before the target metric exists
Before naming an algorithm, write down the label, the prediction time, and the action that changes when the score crosses a threshold. If you cannot say what decision the output drives, any modelling choice is guesswork dressed up as method.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you determine if a sample size is statistically significant bef…
How do you determine if a sample size is statistically significant before making a provider recommendation to a user?
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Say what the estimate is of, and over what population it generalises.
- Write down the assumption the method needs before you use the method.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
How do you measure and rank doctor performance when you have highly va…
How do you measure and rank doctor performance when you have highly varying sample sizes of patient outcomes for each doctor?
Approach
- Write down the assumption the method needs before you use the method.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Sessionise transfer chains into episodes, then count readmissions
encounter has encounter_id, patient_id, facility_id, encounter_type, admit_ts, discharge_ts, admission_type, discharge_disposition, is_planned_admission, principal_diagnosis_code. Chain acute-to-acute transfers into episodes of care: an inpatient encounter discharged as transfer_acute, followed by another inpatient encounter for the same patient at a different facility admitting within 24 hours of that discharge, belongs to the same episode. Then compute the 30-day unplanned readmission rate over episodes, excluding index episodes that are planned, that end in expired, hospice or against medical advice, or that are still open. Return encounter_id to episode_id, plus the rate.
Approach
- Restrict to inpatient encounters for the chaining step and sort by patient_id, admit_ts. Observation and emergency encounters are not acute inpatient stays and pooling them changes both the chain and the denominator.
- Build a boolean continues flag: the previous row is the same patient, its discharge_disposition is transfer_acute, its facility_id differs from this row's, and admit_ts minus previous discharge_ts is between 0 and 24 hours. Episode id = cumsum of the negation of that flag. This is the vectorised form of a sessionisation loop and is why the pattern generalises to any event stream.
- Collapse to episode grain taking first admit_ts, last discharge_ts, last discharge_disposition, and any() of is_planned_admission. The episode inherits its exit state, not its entry state, which is why a transfer chain ending at home is one home discharge rather than two transfers.
- Apply index exclusions at episode grain. Dropping episodes that end in death is not optional: a dead patient cannot be readmitted, so leaving them in inflates the denominator and depresses the rate for exactly the panels caring for the sickest members.
- For each surviving index episode, find the next unplanned acute inpatient episode for that patient admitting within 30 days of the index discharge_ts. Count events at episode grain, since one patient can contribute several index episodes. Report the raw rate and say plainly that it is not comparable across panels without risk adjustment, which is what the observed-over-expected form exists for.
Worked solution 45 min
- Filter to encounter_type 'inpatient', sort by patient_id and admit_ts, and shift discharge_ts, discharge_disposition, facility_id and patient_id by one row.
- Compute continues as same patient AND prior disposition is transfer_acute AND facility differs AND 0 <= (admit_ts - prior discharge_ts) <= 24 hours; episode_id = (~continues).cumsum().
- Aggregate to episodes: first admit_ts, last discharge_ts, last discharge_disposition, any planned flag, patient_id.
- Drop index-ineligible episodes: planned, disposition in expired, hospice or ama, and null discharge_ts.
- For each eligible index episode, search the same patient's later episodes for an unplanned admit_ts within 30 days of index discharge_ts; a merge_asof on patient with a forward direction and a 30-day tolerance does this without a nested loop.
- Rate = index episodes with an event divided by eligible index episodes; return the encounter-to-episode map alongside it.
Follow-up
- discharge_ts is null on an open stay in the middle of a chain. What does your flag do, and what should it do?
- A patient is transferred out and then back to the originating facility within 24 hours. Should that chain, and does your facility_id condition handle it?
- Is a readmission itself eligible to serve as a later index episode? Say what you chose and what the choice does to the rate.
Given a database with two tables (one for actors/patients and one for …
Given a database with two tables (one for actors/patients and one for conditions/procedures), write a query to calculate how often a specific procedure occurs within a group of patients diagnosed with a particular condition.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a query using window functions to identify the most recent medic…
Write a query using window functions to identify the most recent medical event for each patient in a longitudinal dataset.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Compute allowed PMPM by incurred month without span fan-out
Report allowed PMPM by incurred month and product_type for 2025 from medical_claim_line (member_id, claim_id, claim_version, frequency_code, service_start_date, allowed_amount, claim_status), pharmacy_claim (member_id, fill_date, allowed_amount, reversal_flag, reversed_fill_id), member_enrollment (span_id, member_id, product_type, effective_date, termination_date) and member_risk_score (member_id, score_year, prospective_risk_score), which holds at most one row per member per score year. The numerator is allowed_amount on surviving claim versions plus unreversed fills, attributed to the product_type in force on the service or fill date. The denominator is member-months. Also return a risk-adjusted PMPM for each cell: allowed PMPM divided by the member-month-weighted mean prospective_risk_score for score_year 2025 over the members making up that cell's denominator. A colleague joined claim lines to member_enrollment on member_id alone and the annual total came back 1.4 times the finance figure. Diagnose that, then produce the correct monthly series.
Approach
- Name the defect precisely. member_enrollment is one row per span, so joining on member_id alone is one-to-many and duplicates every claim line once per span that member ever held. The inflation factor is the spend-weighted mean span count per member, which is why 1.4 is plausible and why it is not a constant you can divide back out.
- Make the join functional: member_id AND service date between effective_date and COALESCE(termination_date, DATE '9999-12-31'). That yields at most one span per line unless spans genuinely overlap, in which case pick a documented precedence rule, most recently effective wins being the usual one, and prove uniqueness with a duplicate check rather than asserting it.
- Aggregate numerator and denominator separately and join them on (month, product_type). Computing both sides across one join is the second route to inflation, because the member-month denominator fans out over claim lines just as easily as the dollars fan out over spans.
- Clean each numerator source on its own terms: surviving claim versions with voids removed and claim_status 'paid' on the medical side, reversal rows and their originals removed on the pharmacy side.
- Apply completion before anyone reads a trend. State the paid-through date and either withhold the three most recent incurred months or apply completion factors, since the lag differs by service type and the incomplete tail reads as a cost improvement and a utilisation drop at once.
- Risk-adjust last, and weight the score the way the denominator is weighted. The divisor for a cell is SUM(member_months * prospective_risk_score) / SUM(member_months) over that cell's members, not a plain AVG over distinct members: a member covered two months must not count the same as one covered twelve. member_risk_score is member-level, so join it after the member-month grain is fixed, or it becomes a second fan-out.
- State the rule for members with no 2025 member_risk_score row: keep them in member_months, leave them out of the weighted mean, and publish the scored share of member-months beside the adjusted series. A thin score table moves the adjusted number while the raw one sits still, and without that share nobody can tell the two apart.
Worked solution 45 min
- CTE med: surviving claim version per claim_id, voids dropped, claim_status 'paid', service_start_date in 2025.
- CTE rx: fills with reversal_flag FALSE and not referenced by any reversed_fill_id, via NOT EXISTS, fill_date in 2025.
- Attribute each numerator row to one span with the date-bounded join, after asserting that row counts before and after the join are equal.
- CTE denom: member-months by month, product_type and member_id from spans alone, with fractional partial months and overlapping spans collapsed first.
- Join numerator to denominator aggregated on (incurred_month, product_type) and divide for allowed_pmpm.
- Join denom to member_risk_score on member_id for score_year 2025, compute the cell divisor as SUM(member_months * prospective_risk_score) / SUM(member_months) over scored members, carry the scored share of member-months, and divide allowed_pmpm by the divisor for risk_adjusted_pmpm.
- Withhold the three most recent incurred months and label the paid-through date on every row.
Follow-up
- A member holds two overlapping spans under different product_types on the service date. Which one gets the dollar, and how do you keep the member-month denominator consistent with that choice?
- Write the check that proves the span join is one-to-one rather than assuming it.
- Finance reports on paid_date and you reported on service_start_date. Reconcile one month between the two bases.
How would you design a recommendation engine that balances doctor qual…
How would you design a recommendation engine that balances doctor quality, cost-efficiency, and geographical availability for a user?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design an A/B testing framework to measure the success o…
How would you design an A/B testing framework to measure the success of a new member engagement campaign in our mobile app?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
Describe a situation where you had to solve an ambiguous problem under…
Describe a situation where you had to solve an ambiguous problem under a tight deadline. How did you prioritize your tasks?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
What external or alternative data sources would you suggest integratin…
What external or alternative data sources would you suggest integrating to supplement doctor performance rankings when internal claims data is sparse?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Redesign an adherence metric an outreach team cannot quietly inflate
An outreach team is paid on the share of members reaching proportion of days covered of 0.80 or higher in a therapeutic class, computed from pharmacy_claim (member_id, therapeutic_class_code, fill_date, days_supply, refill_number, reversal_flag, reversed_fill_id). The denominator is members with at least two fills in the class and at least 91 days of follow-up, over a 12-month window. Name three ways the team raises this number without any member taking more medication, then redesign the metric and its guardrails so each lever is closed or visible.
Approach
- Attack the denominator first, because it is the cheapest lever. A member who fills once and abandons never enters the denominator at all, so the worst outcome in the class is invisible, and any effort spent on first-fill abandonment lowers the reported share by moving those members in. The incentive runs backwards.
- Attack the window arithmetic second. With the denominator running from first fill to window end, a member whose first fill lands late in the year has a short denominator that is trivially easy to cover, so steering new starts to the last quarter raises the metric mechanically. Report by index quarter or fix the window and require the index fill in the first quarter.
- Attack the threshold third. Days of supply dispensed late in the window count only up to the window end, but a fill placed 60 days before close, on a member whose last on-hand day has already passed, adds up to 60 covered days, about 0.16 of a 365-day denominator. That alone moves a member from 0.77 to 0.93 with no change in what anyone swallows.
- Close the dispensing-side hole: a fill that is billed and reversed is not exposure. Exclude both rows of a reversed pair via reversal_flag and reversed_fill_id, and collapse overlapping fills by pushing each interval start to the later of fill_date and the prior interval's end, so early refills shift coverage forward rather than stacking.
- Redesign the scorecard as a pair rather than a single number: the threshold share stays as the headline, and mean PDC plus the share at 0.90 or higher travel with it, because a distribution cannot be bunched at a threshold without showing it.
- Add the guardrails that make each lever visible: first-fill abandonment rate (fills with refill_number 0 never followed by a second fill, plus reversed first fills), the share of covered days originating from fills in the final 60 days of the window, and the distribution of PDC in the bands just below and just above 0.80.
Worked solution 30 min
- Write the exposure-interval construction explicitly: drop reversed pairs, sort fills per member and class by fill_date, set each interval start to the greater of fill_date and the prior interval end, and truncate the final interval at the window end.
- Write the denominator in two variants, the incumbent at-least-two-fills rule and an index-fill rule requiring the first fill in the first quarter with at least 273 days of follow-up, and state what each includes that the other does not.
- Enumerate the three gaming levers with the arithmetic that makes each one work, including the 60-day end-of-window fill computed against a 365-day denominator.
- Specify the headline metric plus the two companions (mean PDC, share at 0.90 or higher) and the three guardrails, each with numerator, denominator and window.
- Write one sentence stating that dispensing is a proxy for taking and that the proxy overstates adherence in the direction of anyone who fills and does not take.
Follow-up
- The share at 0.80 or higher rose four points and mean PDC rose 0.4 points. What is your reading, and what do you look at next?
- Would you pay on 0.80 at all, or on the continuous measure? Say what each choice costs you.
- Possession is not ingestion. Name a measurable consequence in another table that would corroborate a real adherence improvement, and state its own weakness.
Pooled persistency fell but no one lost coverage
Pooled 12-month enrollment persistency fell from 0.78 to 0.66 for the cohort anchored in a single month. member_enrollment carries span_id, member_id, plan_id, product_type, effective_date, termination_date, termination_reason_code, dual_eligible_flag and region_code. The metric allows at most one gap of 45 days or fewer. Product mix in that anchor month was not stable. Determine how much of the fall is real disenrollment rather than span bookkeeping or mix, and deliver a persistency series a contract owner could act on.
Approach
- Inspect the query's unit of analysis first. Spans are not members, and a benefit or plan change closes one span and opens the next on the same day, so any logic that treats a non-null termination_date as loss of coverage counts an administrative event as churn.
- Rebuild member-level coverage by unioning all spans per member and merging adjacent or overlapping ones, then apply the 45-day gap rule to the merged intervals rather than per span.
- Cut the terminations by termination_reason_code. plan_change and eligibility_redetermination have entirely different meanings for a contract, and death is not churn you can address.
- Split by product_type and stop pooling. Commercial open enrollment and managed medicaid redetermination produce structurally different churn on different calendars, and the pooled average of the two is a number that describes nobody.
- Compare the anchor month to the same month in the prior year rather than to the adjacent month, because enrollment is strongly seasonal and an anchor month sitting on an open-enrollment or redetermination boundary is not comparable to its neighbour.
- Apply the Kitagawa decomposition across product_type to separate how much of the 12 point fall is mix, meaning the anchor cohort simply contained more of a high-churn product, from how much is within-product deterioration.
Follow-up
- Why 45 days, and what would change if the allowable gap were 30 or 60?
- A member terminates and re-enrolls 90 days later under a different plan_id. Are they retained? Defend your answer against the opposite one.
- What guardrail would catch someone improving persistency by declining to enroll members likely to churn?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.
How do you handle `NULL` values when performing left joins on large he…
How do you handle NULL values when performing left joins on large healthcare claims tables to ensure your metrics remain accurate?
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that project again?
Sequence three urgent requests with one analyst-week available
Three requests land on Monday and you have one week. An actuarial team needs incurred-claims completion factors restated before a filing deadline on Thursday. A clinical programme owner wants a deterioration model refreshed because its calibration has drifted in one region. A trial operations team wants site enrolment forecasts for a portfolio review in two weeks. Each requester believes theirs is blocking. Deliverable: your sequence with the reasoning, the message you send to whoever is deprioritised, and the smaller artefact you hand each of the two you cannot fully serve.
Approach
- The probe is whether you prioritise on consequence and reversibility rather than on who asked loudest or most recently.
- Classify each request by what happens if it slips. A regulatory or contractual deadline is irreversible on its date, a drifting model is causing harm every day it keeps running, and a portfolio review can absorb a provisional number. That ordering is defensible to all three requesters because it does not depend on your preferences.
- Take the deadline-bound work first, but scope it to the minimum defensible output, because completion factors feeding a filing carry a different error tolerance than a slide.
- Do not let the drifting model simply wait. Quantify the harm cheaply by comparing calibration in the affected region against the rest, and if it is materially miscalibrated propose flagging or suppressing its output for that region within the hour rather than at the end of a refresh.
- Give each deprioritised requester something real: a provisional forecast with its uncertainty and a refresh date, or a diagnostic that tells them whether their problem is urgent. Say no explicitly with a date rather than going quiet, because silence is what produces escalation.
Follow-up
- The programme owner escalates to your manager. What do you want your manager to be able to say?
- Midweek the actuarial work needs two more full days than you estimated. What gives?
- How does your answer change if the drifting model drives a clinical outreach list rather than a report?
Scope a one-line request for a readmission rate
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
Approach
- Read the probe: this tests whether you convert ambiguity into a specification without either stalling or guessing silently. Both failure modes are common, and the silent guess is worse because nobody can see it.
- Rank your questions by leverage on the number rather than by curiosity. Whether the numerator is all-cause or unplanned, and whether the clock starts at discharge_ts, move the result far more than a tie-break rule on same-day returns.
- Ask about the comparison first. A number going to a committee will be compared to something, and whether that something is last quarter, another panel, or a national figure decides whether you owe them a raw rate, a risk-adjusted rate, or an observed-over-expected ratio.
- Commit to a default in writing so silence does not block you: unplanned acute inpatient readmission within 30 days of discharge_ts, index stays excluding planned admissions, acute-to-acute transfers, discharges against medical advice, and dispositions of expired or hospice, denominator of eligible index discharges, paid-through date stated.
- Write the caveat as a restriction on use rather than a hedge: name the one comparison the figure supports and the one it does not.
Follow-up
- They reply that they want it 'the way the benchmark does it'. What do you ask next?
- The committee wants it split by attending provider. What changes in the specification, and what do you refuse to show?
- How does your answer change if the meeting is tomorrow rather than Friday?
- 01
How do you handle `NULL` values when performing left joins on large healthcare claims tables to ensure your metrics remain accurate?
- 02
Three requests land on Monday and you have one week. An actuarial team needs incurred-claims completion factors restated before a filing deadline on Thursday. A clinical programme owner wants a deterioration model refreshed because its calibration has drifted in one region. A trial operations team wants site enrolment forecasts for a portfolio review in two weeks. Each requester believes theirs is blocking. Deliverable: your sequence with the reasoning, the message you send to whoever is deprioritised, and the smaller artefact you hand each of the two you cannot fully serve.
- 03
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
Is this an official Garner health interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Garner health. Rounds and questions reflect what candidates have reported, not a process Garner health has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the interview process compared to other data science roles?
The process is highly technical but balanced. While the initial statistical screen is relatively straightforward, the subsequent SQL challenges and take-home/case study assessments are rigorous. They evaluate not just your mathematical knowledge, but your ability to write clean, production-grade code and think strategically about the product.
PracHub interview research ↗What is the hybrid work policy for this position?
For roles based in the New York City office, Garner Health operates on a hybrid model. You must be willing to work in the office 3 days per week, specifically on Tuesday, Wednesday, and Thursday.
PracHub interview research ↗How important is prior healthcare experience for this role?
While prior experience with healthcare data (like claims or medical codes) is a strong plus, it is not a strict requirement. Garner Health values strong first-principles thinking, statistical rigor, and engineering excellence. They are fully prepared to help talented, mission-driven data scientists build domain expertise on the job.
PracHub interview research ↗What is the company culture like, and how is it evaluated in the interviews?
Garner Health is a fast-growing, mission-driven startup that operates with intense urgency and a high level of individual accountability. In your behavioral interviews, the team will look for candidates who are comfortable with direct, authentic feedback and who thrive in collaborative, high-velocity environments.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22