As a Data Scientist at Healthfirst (New York), you play a pivotal role in harnessing data to drive strategic decisions and improve healthcare outcomes. Your expertise in analyzing complex datasets will directly influence the development of health programs, enhance operational efficiency, and contribute to patient-centric solutions. This role is vital not only for the growth of Healthfirst but also for the overall well-being of the communities it serves.
In this position, you will collaborate with cross-functional teams, including product managers, engineers, and healthcare professionals, to tackle real-world problems. You will be involved in projects that analyze patient data, optimize care pathways, and develop predictive models that can foresee healthcare trends. Your work will empower the organization to make informed decisions, ultimately leading to improved patient experiences and better health results.
Expect to engage with a variety of data sources and analytical techniques, making this role not only critical but also intellectually stimulating. You will encounter challenges that require innovative thinking and a deep understanding of both data science and healthcare, setting the stage for a rewarding career at.
Initial Screening
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Assessments
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Behavioral Interviews
reportedRounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.
What to demonstrate
- Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
- What you changed afterwards, stated as a check you now run rather than a lesson you now believe
- Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you
How to prepare
- Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
- Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
- Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
PracHub editorial advice for the preparation topics above.
Treating clinical measurements as missing at random
A lab result, a vital sign, or a screening exists because someone ordered it, and ordering tracks suspicion of disease, visit frequency, and site workflow. Imputing the mean or dropping incomplete rows biases the population estimate and can flip the sign of an association, because the untested are systematically healthier or systematically disengaged. The presence indicator is often more predictive than the value, which is a warning sign rather than a feature win: a model that learns test ordering will not transfer to a site with different protocols.
Reading the most recent months of a claims-based series as real
Claims incur before they are reported and paid, so recent incurred months are systematically undercounted until runout completes. The lag is not uniform: pharmacy adjudicates in days, professional claims in weeks, inpatient facility claims in months. That means recent data is both too low and mix-shifted toward cheap services, which reads as a cost improvement and a utilisation drop at once. The fix is to hold the last three incurred months back or apply completion factors, and to state the paid-through date on every chart.
Crediting a treatment for regression to the mean
Selecting a group because it is extreme (lowest-engagement users, accounts having their worst month, the bottom decile of a score) moves that group's expected next-period value back toward the average even under no treatment, by exactly as much as the selecting measure is imperfectly correlated with its own later value. Compare against units that met the same selection rule and went untreated, or use two pre-periods so the bounce-back is visible before the intervention starts. A pre-post number on a group chosen for being extreme measures the selection rule, not the treatment.
Reading experiment results before checking the arm split
Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
What metrics do you use to evaluate model performance?
What metrics do you use to evaluate model performance?
Approach
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Simulate how ignoring clustering inflates the false positive rate
Simulate a provider-level comparison with no true effect. 40 providers, 60 members each, a continuous outcome with total variance 1 and intraclass correlation 0.02. Randomise providers 20 to 20, not members. Over 5,000 replications, report the share of replications where a two-sample t-test run on all 2,400 individual observations, ignoring provider, returns p below 0.05. Report the same share when the unit of analysis is the 40 provider means. Then state how the first number follows from the design effect 1 + (m - 1) x ICC.
Approach
- Decompose the variance explicitly rather than tuning it: provider variance = ICC = 0.02, residual variance = 1 - ICC = 0.98. Draw u_j once per provider and e_ij per member, so the outcome is u_j + e_ij and the correlation between two members of one provider is 0.02 by construction.
- Randomise at the provider level. That is the entire mechanism: the treatment indicator is constant within a provider, so in any single replicate the provider intercepts are confounded with the arm, and the naive test reads that confounding as signal.
- Vectorise over replications with a (reps, providers, members) array. 5,000 x 2,400 normals is trivial in numpy, and a Python loop is what tempts people to cut replications to 500 and report a number with Monte Carlo noise of plus or minus 0.017.
- Predict the answer before running it: design effect = 1 + (60 - 1) x 0.02 = 2.18, naive standard errors are too small by sqrt(2.18) = 1.476, so the rejection rate at nominal 0.05 is about 2 x (1 - Phi(1.96 / 1.476)) = 0.18. Agreement between prediction and simulation is what makes the result a demonstration rather than an anecdote.
- Run the provider-mean analysis as the control. Recovering 0.05 there proves the generator is correct and isolates the failure to the analysis unit.
Follow-up
- How many providers would you need for 80 percent power on a 0.2 SD difference with 60 members each and this ICC?
- Cluster sizes are unequal in reality. The design effect approximation becomes 1 + ((1 + CV^2) x mbar - 1) x ICC. Which direction does that move your sample size and why?
- The intervention cannot be withheld from any provider. Sketch a design that still yields a defensible estimate.
Proportion of days covered with shifted, truncated refill intervals
pharmacy_claim has fill_id, member_id, therapeutic_class_code, fill_date, days_supply, reversal_flag, reversed_fill_id. For one therapeutic class and a fixed 12-month window, compute proportion of days covered per member: distinct days on which a dispensed days_supply covers the day, divided by days from the member's first in-window fill through the window end. An early refill shifts coverage forward rather than stacking, and coverage is truncated at the window end. Drop both rows of every reversed pair. Return member_id, pdc, and the share at pdc >= 0.80 among members with at least 2 fills and at least 91 days of follow-up.
Approach
- Remove reversals as pairs first. Drop every row with reversal_flag true, and also drop the fill_ids those rows point at through reversed_fill_id. Dropping only the flagged row leaves a dispense that was never collected in the exposure.
- Walk fills per member in fill_date order carrying a cursor: start = max(fill_date, previous_end + 1 day), end = start + days_supply - 1. The shift is path dependent, so a plain cumsum over days_supply does not reproduce it. Use itertools.accumulate or a per-member loop over numpy arrays, not a row-wise apply over the whole frame.
- Truncate the last interval at the window end before measuring. Without truncation a 90-day fill dispensed on the final day pushes the covered-day count past the denominator and PDC above 1.0.
- Use the denominator the definition states: first in-window fill_date through window end, inclusive. Not a flat 365, and not first fill to last fill, which is a different metric that rewards early discontinuation.
- Apply the eligibility filter before computing the >= 0.80 share, and report the size of that denominator next to the share. A share without its denominator is not reviewable.
Worked solution 40 min
- Filter to the therapeutic class and the window, then remove reversed pairs by dropping flagged rows and the fill_ids in reversed_fill_id.
- Sort by member_id, fill_date. Per member, accumulate start = max(fill_date, prior_end + 1 day) and end = start + days_supply - 1.
- Clip every interval's end at the window end and drop intervals whose start is past it.
- Covered days per member = sum of (end - start + 1) over the shifted intervals, which are disjoint by construction.
- Denominator = (window_end - first_fill_date).days + 1. Divide, then filter to members with >= 2 fills and denominator >= 91 and compute the share at or above 0.80.
Follow-up
- How does PDC differ from medication possession ratio, and which of the two can exceed 1.0?
- A member switches to a different ingredient inside the same therapeutic class mid-window. Should the intervals chain, and what does that do to the class-level number?
- Members who die or lose coverage mid-window get a short denominator and often a high PDC. If you then compare mortality by adherence category, what bias have you built in and how do you remove it?
Count thirty-day unplanned readmissions without counting transfers
encounter holds encounter_id, patient_id, facility_id, encounter_type, admit_ts, discharge_ts, admission_type, discharge_disposition, is_planned_admission and index_encounter_id. For inpatient stays discharged between 2025-01-01 and 2025-11-30, return the 30-day unplanned readmission rate overall and by facility_id. Exclude an index stay when is_planned_admission is TRUE or discharge_disposition is transfer_acute, expired, hospice or ama. A qualifying readmission is any unplanned inpatient admission for the same patient with admit_ts after the index discharge_ts and within 30 days of it. One stay can be both a readmission and the next index stay.
Approach
- Decide between LEAD and EXISTS explicitly and say why. LEAD over the patient's admission stream gives the literal next stay, so an intervening planned admission hides a later unplanned one. The measure asks whether any unplanned admission fell in the window, which is an EXISTS or a self-join with a range predicate.
- Exclude transfers at the index end before anything else. An acute-to-acute transfer closes one encounter and opens another at a different facility_id within hours, so without the transfer_acute exclusion the facilities that move the sickest patients out look the worst.
- Apply the index exclusions as a set, including stays with NULL discharge_ts, which are still open and have no measurable window.
- Stop the denominator early enough that every index stay had a full 30 days of observation plus encounter lag. That is why the window ends 2025-11-30 rather than at year end.
- Compute the rate at stay level, not patient level, then compare the raw facility rates against an observed-over-expected framing before anyone publishes a ranking.
Worked solution 30 min
- CTE ip: inpatient encounters with a non-null discharge_ts.
- CTE index_stays: ip rows with discharge_ts in the window, is_planned_admission FALSE, and discharge_disposition not in the four excluded values.
- Flag each index stay with EXISTS against ip for the same patient_id where is_planned_admission is FALSE, encounter_id differs, admit_ts > index.discharge_ts and admit_ts <= index.discharge_ts + 30 days.
- Aggregate numerator and denominator overall and grouped by facility_id.
- Re-run with the transfer exclusion removed and record how much the rate moves.
Follow-up
- A patient has three stays ten days apart. How many index stays and how many readmissions, and why does the answer depend on the stated rule?
- Which comparison would you publish across facilities, raw rate or observed over expected, and what does risk adjustment still fail to fix?
- index_encounter_id already exists on the table. Would you trust it, and how would you audit it against your own logic?
Find inpatient encounters with no resulted lab during the stay
encounter holds encounter_id, patient_id, encounter_type, admit_ts, discharge_ts and facility_id. lab_result holds result_id, order_id, patient_id, encounter_id (NULL for outpatient standing orders), loinc_code, specimen_collected_ts, resulted_ts and result_status. Return every inpatient encounter discharged in 2025 with no final or corrected lab_result linked to it. A colleague wrote WHERE encounter_id NOT IN (SELECT encounter_id FROM lab_result) and got zero rows back. Say why, write the correct query, then say what changes when labs must instead be matched on patient_id with the specimen collected between admit_ts and discharge_ts.
Approach
- Name the mechanism rather than the symptom. x NOT IN (subquery) expands to x <> y1 AND x <> y2 AND ..., and a single NULL y makes that chain UNKNOWN, never TRUE. Since lab_result.encounter_id is nullable by design for standing orders, this query can only ever return zero rows.
- Rewrite as NOT EXISTS, or LEFT JOIN with an IS NULL filter. Both treat a NULL key as no match rather than as unknown. Adding IS NOT NULL inside the NOT IN subquery also works, but NOT EXISTS is the habit that survives someone making another column nullable later.
- Put the result_status predicate inside the correlated subquery or the ON clause, not in an outer WHERE. Outside, it turns the anti-join back into an inner join and an encounter whose only labs were cancelled disappears instead of qualifying.
- For the timestamp variant, match on patient_id with specimen_collected_ts inside the stay rather than resulted_ts. A specimen drawn an hour before discharge can result the next day, and anchoring on resulted_ts would wrongly call that stay lab-free.
- Restrict to encounter_type 'inpatient' and discharge_ts in 2025, and report open encounters with a NULL discharge_ts as a separate count instead of letting the filter swallow them.
Follow-up
- Write the LEFT JOIN form and say exactly where the result_status predicate must sit for the two forms to agree.
- How would you separate encounters whose only labs were cancelled from encounters with no lab rows at all?
- Some encounters carry member_id NULL because the person index did not match. What does that do to a payer-side version of this measure?
If tasked with reducing patient readmission rates, what approach would…
If tasked with reducing patient readmission rates, what approach would you take?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Walk me through your thought process for a project where you had limit…
Walk me through your thought process for a project where you had limited data.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How do you prioritize tasks when working on multiple projects?
How do you prioritize tasks when working on multiple projects?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
Given a dataset with patient admissions, how would you identify trends…
Given a dataset with patient admissions, how would you identify trends in hospital utilization?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design an A/B test for a new health initiative?
How would you design an A/B test for a new health initiative?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Explain the differences between supervised and unsupervised learning.
Explain the differences between supervised and unsupervised learning.
Approach
- Clarify what is being asked and what a complete answer would contain.
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Decompose risk-adjusted cost per member into owned leaf metrics
Build a metric tree, at most four levels deep, from risk-adjusted allowed PMPM down to leaves an owner can move. Data: medical_claim_line (allowed_amount, units, procedure_code, place_of_service_code, service_start_date, claim_version, frequency_code), pharmacy_claim (allowed_amount, days_supply, is_generic, formulary_tier, reversal_flag), member_enrollment (coverage spans, prospective_risk_score). Give the arithmetic identity each split rests on, the leading indicator at each leaf, and identify the one node where the arithmetic is exact but the attribution is not.
Approach
- State the base identity before drawing anything. Allowed PMPM equals total allowed over member-months, and it decomposes exactly by service category c as the sum over c of (units_c / member-months) times (allowed_c / units_c). Utilization per member-month and allowed per unit, summed across categories, so category mix is inside the sum rather than a separate hand-wave.
- Give the change decomposition too, because the tree will be read on deltas: the change in PMPM equals the sum over categories of (change in utilization times base price) plus (base utilization times change in price) plus (change in utilization times change in price). The third term is the interaction, it is not noise, and an analysis that drops it will not reconcile.
- Build the medical branch to leaves with owners: inpatient admits per 1,000 member-months and allowed per admit (with DRG case mix as the sub-split), ambulatory visits per 1,000 and allowed per visit split by place_of_service_code, and emergency visits per 1,000 and allowed per visit. Prepare the data with version dedup to the surviving claim_version before a single dollar is summed, since frequency_code 7 replaces and 8 voids.
- Build the pharmacy branch on 30-day equivalents: scripts_30 equals the sum of days_supply over 30 on unreversed fills, so pharmacy PMPM equals (scripts_30 per member-month) times (allowed per 30-day script). Split the price term exactly as g times the generic price plus (1 minus g) times the brand price, where g is the generic dispensing rate, which makes generic substitution and unit price separable leaves with different owners.
- Attach one leading indicator per leaf: prior-authorization approvals and elective surgical bookings ahead of inpatient admits, contract effective dates ahead of allowed per unit, site-of-service steering ahead of the ambulatory place_of_service mix, new-to-brand starts and formulary tier placement ahead of pharmacy price.
- Name the node where arithmetic is exact but attribution is not: the risk-adjustment denominator. Dividing by the population mean prospective_risk_score is exact arithmetic, but the score is built from coded diagnoses, so more complete coding raises the denominator and lowers risk-adjusted PMPM with no change in care delivered. Carry raw PMPM, the mean risk score trend and a coding-intensity measure such as diagnoses per member-year beside it, permanently.
Worked solution 45 min
- Write the base identity and the three-term change decomposition on paper first, and confirm the terms sum to the total change on a two-category toy example before touching real data.
- Specify the preparation rules that precede any aggregation: dedup medical_claim_line to the surviving version per claim_id using claim_version and frequency_code, drop both rows of every reversed pharmacy pair, and build member-months from coverage spans with partial months counted fractionally.
- Draw the medical branch and the pharmacy branch to leaf level, writing the identity beneath each split rather than only the metric name.
- Write the generic-dispensing split as the exact price identity and confirm it reproduces the observed allowed per 30-day script.
- Attach one leading indicator and one named owning function to each leaf.
- Write the caveat on the risk-adjustment node, including the two companion series that must be published with it.
Follow-up
- Utilization is flat and allowed per unit is flat, yet PMPM rose 4 percent. Name every place the increase can be hiding.
- Your member-month denominator came from coverage spans that overlap across products for some members. What does that do to PMPM, and how do you decide whether to deduplicate?
- Which leaf would you refuse to put on a team's scorecard, and why?
Glycemic control share jumped when a feed connected
The share of members with a controlled glycemic result moved from 0.62 to 0.81 the month a second result feed was connected. lab_result carries result_id, order_id, patient_id, encounter_id, loinc_code, analyte_name, value_numeric, units, reference_low, reference_high, abnormal_flag, result_status and supersedes_result_id. The metric is members whose latest result in the year is below the control threshold, divided by members with any result in the year. Decide whether control actually improved, and specify the metric so results from both feeds can be pooled.
Approach
- Split numerator and denominator. The denominator is members with any result, so a feed that tests more people changes the metric without changing anyone's health. Count denominator members before and after and see how much of the 19 point move is denominator growth.
- Recompute on a fixed cohort: members who had a qualifying result in both the before and after periods. If the share is flat on the fixed cohort, the movement is who got tested, not what their results were.
- Deduplicate to the latest non-cancelled row per (order_id, loinc_code), respecting supersedes_result_id, because a corrected result arrives as a new row and counting both gives one person two values.
- Audit units and loinc_code mix by feed. Two unit conventions coexist for this analyte, percent and mmol/mol, related by mmol/mol equal to 10.929 times (percent minus 2.15), so 7.0 percent is 53 mmol/mol. Comparing mmol/mol values to a percent threshold would classify them all as uncontrolled and depress the share, so a rise cannot be explained by that mismatch, which rules it out rather than in.
- State plainly that the tested population is not the covered population. Testing is ordered when someone suspects a problem or when a site runs standing orders, so the mean among tested members is not the population mean and cannot be made into one by imputation.
- Respecify: numerator and denominator both restricted to members with a stated continuous-enrolment requirement, one deduplication rule, one unit after conversion, one threshold, and a separately reported testing rate so the denominator's movement is visible.
Follow-up
- Someone proposes imputing the mean result for untested members. What happens to the estimate and why?
- A model you are building gains AUC from a feature that marks whether the test exists. Why is that a warning rather than a win?
- How would you report this metric so a clinical committee can see improvement separately from coverage of testing?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.
Can you give an example of a challenging team dynamic you faced and ho…
Can you give an example of a challenging team dynamic you faced and how you managed it?
Approach
- Quantify the outcome, including what you would not claim credit for.
- State the situation in two sentences and spend the rest on your reasoning.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Describe a time when you used a specific algorithm to solve a problem.
Describe a time when you used a specific algorithm to solve a problem.
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Disagree with a product manager about an adherence nudge
A product manager proposes shipping a refill-reminder feature to the whole member base, citing an internal analysis: members with 12-month proportion of days covered at or above 0.80 had 31 percent fewer admissions than members below it. The roadmap date is three weeks out and engineering is scoped. You believe the comparison builds survival and continuous coverage into the exposure definition, and is separately confounded by the kind of member who refills on time. Deliverable: what you say in the room, what you send afterwards, and the smallest study you would accept as sufficient to ship.
Approach
- The probe is whether you can be specific and non-territorial at once. A disagreement that arrives as 'the analysis is flawed' costs you the room; one that names a mechanism and offers a cheaper path keeps it.
- Name the two defects separately, because they have different fixes. Classifying a member as adherent over 12 months requires them to stay alive and stay covered for those 12 months, and that guaranteed event-free interval is assigned to the adherent group. Separately, members who refill on schedule differ from those who do not in ways claims never record.
- Show rather than assert. Rerun the same comparison as a landmark analysis: classify adherence over days 1 to 90, start follow-up at day 91 for everyone still enrolled and event-free, and report how much of the 31 percent survives.
- Separate the association claim from the intervention claim out loud. Nothing in the analysis speaks to whether a reminder changes refill behaviour, which is the actual product question and is separately testable.
- Offer the shippable path: a randomised holdout on a subset with proportion of days covered as the primary endpoint, powered on the plausible effect of a reminder rather than on the 31 percent, with a stated read date.
Follow-up
- The landmark rerun still shows a 20 percent difference. Do you ship?
- The product manager says a holdout delays launch. How do you cost that argument?
- What primary endpoint do you choose, and why not admissions?
- 01
Can you give an example of a challenging team dynamic you faced and how you managed it?
- 02
Describe a time when you used a specific algorithm to solve a problem.
- 03
A product manager proposes shipping a refill-reminder feature to the whole member base, citing an internal analysis: members with 12-month proportion of days covered at or above 0.80 had 31 percent fewer admissions than members below it. The roadmap date is three weeks out and engineering is scoped. You believe the comparison builds survival and continuous coverage into the exposure definition, and is separately confounded by the kind of member who refills on time. Deliverable: what you say in the room, what you send afterwards, and the smallest study you would accept as sufficient to ship.
Is this an official Healthfirst (New York) interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Healthfirst (New York). Rounds and questions reflect what candidates have reported, not a process Healthfirst (New York) has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the typical difficulty level of interviews for this position?
The interview difficulty for the Data Scientist role at Healthfirst is generally considered average. Candidates should be prepared for a mix of technical and behavioral questions, with emphasis on practical applications of data science.
PracHub interview research ↗How long does the interview process usually take?
The interview timeline can vary, but candidates typically can expect the process to last a few weeks from initial screening to final decisions. It's advisable to stay engaged and proactive in communication during this period.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates often demonstrate a strong blend of technical expertise and soft skills. They effectively communicate their thought processes, show genuine interest in healthcare, and align closely with the organization's mission.
PracHub interview research ↗How does the culture at Healthfirst support professional growth?
Healthfirst fosters a supportive culture that encourages continuous learning and collaboration. Employees have access to professional development resources and are encouraged to engage in team projects that promote innovation.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22