Flatiron Health · Data Scientist
Updated · 2026-09-24

Flatiron Health Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Flatiron Health, your role is pivotal in transforming healthcare data into actionable insights that drive improvements in cancer patient care and research. You will work with complex datasets, leveraging advanced analytics and machine learning to solve critical clinical questions, ultimately impacting the lives of patients and healthcare providers alike. Your work will not only influence internal product development but also contribute to broader healthcare initiatives, making this a highly strategic position within the organization.

Learn the economics of the product category before the loop. Marketplaces, subscription products and ad-supported products turn on different core quantities (match rate and liquidity, retention and churn, fill rate and yield) and fail in different characteristic ways.

Flatiron Health candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Separate censoring from events in time-to-event modelsRisk-adjust outcomes before comparing providers or panelsReconstruct drug exposure intervals from dispensing records

30 min read

Practice 14 Data Scientist prompts
6Company bank questionsSnapshot · Sep 28, 2026 PT
1Candidate experiences ↗Read their reports
14Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Flatiron Health, your role is pivotal in transforming healthcare data into actionable insights that drive improvements in cancer patient care and research. You will work with complex datasets, leveraging advanced analytics and machine learning to solve critical clinical questions, ultimately impacting the lives of patients and healthcare providers alike. Your work will not only influence internal product development but also contribute to broader healthcare initiatives, making this a highly strategic position within the organization.

In this role, you can expect to collaborate with cross-functional teams that include clinicians, engineers, and product managers, focusing on real-world applications of data analysis in oncology. This involves tackling challenging problems that require innovative thinking and a deep understanding of both data science and healthcare. The complexity and scale of the data you will handle provide an exciting opportunity to make a significant difference in the healthcare landscape.

01

Initial Screening

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Assessments

reported

This round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.

What to demonstrate

  • Whether your row counts survive each join, and whether you notice on your own when they do not
  • Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
  • Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
  • Reaching a defensible answer inside the window instead of a refined one after it

How to prepare

  • Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
  • Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
  • Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
PracHub interview research ↗
03

Technical Interviews

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗
04

Behavioral Interviews

reported

Most of the weight in this round sits on the disagreement questions. Data work routinely produces an answer someone senior did not want, and the interviewer is trying to learn what you do in that hour. Both failure modes are common: folding as soon as a director pushes back, and treating the pushback as ignorance to be corrected with a better chart. A strong answer usually contains a specific thing the other person knew that you did not, and describes how you found out whether it changed the conclusion.

What to demonstrate

  • Whether you can state the other side's argument accurately before you explain why you disagreed
  • What you treated as evidence during the disagreement, such as a rerun under their assumption or a holdout check, rather than persuasion technique
  • Whether you distinguish being overruled from being wrong, and can give an example of each

How to prepare

  • Write out one disagreement where you turned out to be wrong, and say what in the data misled you. Candidates prepare the story where they were right, and the follow-up asks for the other one.
  • For your main disagreement story, be ready to say what result would have made you drop your position. If no such result exists, you were not arguing from the data.
  • Practise stating the opposing position out loud in one sentence the stakeholder would accept, then continue the story.
PracHub interview research ↗
05

Final Interviews

reported

A day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.

What to demonstrate

  • Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
  • Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
  • Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method

How to prepare

  • Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
  • Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
  • Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
  • Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub interview research ↗

1 candidate reports. Individual accounts describe a particular role and hiring cycle.

Data Analyst

Flatiron Health Data Analyst Interview Experience — Passed Coding, Cut on the Final Case Round

Technical Screen → OnsiteOutcome: rejected

Let me share a DA interview. I signed an NDA so I won't go into detail. From what I've seen on the forum, it feels like nobody has ever gotten an offer from this one... so I didn't have very high hopes going in, and sure enough, that's how it went. First was the Karat screen: they tested SQL, R, and some statistics knowledge. A week later I got the invite for the first round. Round 1: they asked…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Reading the most recent months of a claims-based series as real

Claims incur before they are reported and paid, so recent incurred months are systematically undercounted until runout completes. The lag is not uniform: pharmacy adjudicates in days, professional claims in weeks, inpatient facility claims in months. That means recent data is both too low and mix-shifted toward cheap services, which reads as a cost improvement and a utilisation drop at once. The fix is to hold the last three incurred months back or apply completion factors, and to state the paid-through date on every chart.

02

Pre-post evaluation on a high-cost or high-risk cohort

Cohorts selected on an extreme value of the outcome regress toward the mean on their own. Members identified as the top 1 percent of spend in one year spend far less the next, whether or not anyone intervenes, because the selecting year captured both chronic severity and one-off events. A pre-post design on such a cohort will report savings every time. A concurrent comparison group selected by the same rule in the same period, or a regression discontinuity at the selection threshold, is the minimum credible design.

03

Writing SQL without stating NULL and tie-breaking behaviour

Before calling a query finished, say what it does with NULLs, ties and empty groups. NOT IN against a subquery containing a single NULL returns no rows at all, and RANK, DENSE_RANK and ROW_NUMBER differ precisely on ties, so name which one the question requires.

04

Solving silently instead of narrating the reasoning

Say which branch you are taking and why you chose it over the alternative, for example checking the denominator first because it changes what the comparison means. A correct answer that arrives with no visible path scores below a rigorous one that needed a hint.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

11 technical prompts3 include a worked solution

What metrics would you use to evaluate a machine learning model's perf…

medium
machine learning and modelling

What metrics would you use to evaluate a machine learning model's performance?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

How would you prioritize features in a predictive model for patient re…

medium
machine learning and modelling

How would you prioritize features in a predictive model for patient readmission?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • How would you choose the decision threshold, and who owns that choice?

Expand overlapping coverage spans into fractional member-months

mediumWorked solution
interval expansiondenominatorsvectorisation

member_enrollment has span_id, member_id, plan_id, product_type, effective_date, termination_date (NULL while active). Both dates are inclusive. Write a function returning one row per calendar month of a stated year with total member-months, where a member contributes covered days in the month divided by days in that month, capped at one month per member even when two spans overlap after a plan change or a retroactive span. Do not expand to one row per member-day: the sample has 400,000 members. Output: month, member_months.

Approach
  1. Resolve NULL termination_date to the reporting end date and say so in a comment. An open span is not an infinite span, and clipping it at the year end is what keeps the denominator finite and auditable.
  2. Merge overlapping and adjacent spans per member before touching months. Sort by member_id and effective_date, carry a running maximum of the end date, and open a new merged group when effective_date exceeds that running max plus one day. Adjacent means a gap of zero days, which a plan change produces constantly.
  3. Cross join the merged spans to the 12 month boundaries rather than to days. Overlap days = (min(span_end, month_end) - max(span_start, month_start)).days + 1, clipped below at 0. That is at most 12 rows per span instead of 365.
  4. Divide overlap days by the number of days in that month, so February and July are weighted correctly, then sum by month.
  5. Validate on a constructed member before trusting the aggregate: a single span covering the whole year must sum to exactly 12.0.
Worked solution 30 min
  1. Fill termination_date nulls with the reporting end date and clip all spans to the reporting year.
  2. Sort by member_id, effective_date, then compute a running max end per member with cummax shifted by one, flag a new group where effective_date > prior_max_end + 1 day, and cumsum the flag to get merged group ids.
  3. Aggregate each group to min start and max end, producing disjoint spans per member.
  4. Cross join merged spans to a 12-row month frame, compute clipped overlap days, divide by days in month.
  5. Group by month and sum.
EXPECTED RESULTTwelve rows. For a member with spans 2026-01-01 to 2026-03-15 and 2026-03-01 to 2026-06-30, March contributes exactly 1.0, not 15/31 plus 31/31. For a member with spans covering January and March only, January is 1.0, February is 0.0, March is 1.0, and the member totals 2.0.
Follow-up
  • Product type changes mid-year. The metric must be reported by product_type. Where does the cap now apply and what breaks?
  • An eligibility file arrives with a retroactive termination that shortens a span you already reported on. How do you restate?
  • Why member-months rather than distinct members, in one sentence, for a director who wants the simpler number?

Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Fix the scope and set a baseline
  • Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
  • Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
  • Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.

Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02One query pattern, written three times
  • Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
  • On the third attempt, write the grain of every CTE as a comment before writing its body.
  • Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.

Deliverable: Three independent versions of the same query plus a note on what changed between them.

Practice prompt ↗Practice prompt ↗
03Only the statistics you will be asked to defend
  • Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
  • Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
  • Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.

Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.

Practice prompt ↗Practice prompt ↗
04One case, and the assumptions holding it up
  • Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
  • Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
  • Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.

Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Your own work, timed
  • Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
  • Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
  • Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.

Deliverable: Two timed narratives with one defensible number in the opening line.

Practice prompt ↗Practice prompt ↗
06The one full rehearsal, in a longer weekend block
  • Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
  • Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
  • Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.

Deliverable: Mock notes naming three failure moments with a specific fix written under each.

Practice prompt ↗Practice prompt ↗
07Taper
  • Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
  • Re-read only your own notes from this week, and open no new material.
  • Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.

Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.

How do you approach disagreements with team members or stakeholders?

medium
behavioural and stakeholder questions

How do you approach disagreements with team members or stakeholders?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Quantify the outcome, including what you would not claim credit for.
  3. Close with what you would do differently, concretely.
Follow-up
  • How did you know the outcome was caused by your change?
  • What would you do differently if you ran that project again?

What motivates you to work in the healthcare sector?

medium
behavioural and stakeholder questions

What motivates you to work in the healthcare sector?

Approach
  1. Close with what you would do differently, concretely.
  2. Quantify the outcome, including what you would not claim credit for.
  3. State the situation in two sentences and spend the rest on your reasoning.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Sequence three urgent requests with one analyst-week available

medium
prioritisationtradeoffsstakeholder communication

Three requests land on Monday and you have one week. An actuarial team needs incurred-claims completion factors restated before a filing deadline on Thursday. A clinical programme owner wants a deterioration model refreshed because its calibration has drifted in one region. A trial operations team wants site enrolment forecasts for a portfolio review in two weeks. Each requester believes theirs is blocking. Deliverable: your sequence with the reasoning, the message you send to whoever is deprioritised, and the smaller artefact you hand each of the two you cannot fully serve.

Approach
  1. The probe is whether you prioritise on consequence and reversibility rather than on who asked loudest or most recently.
  2. Classify each request by what happens if it slips. A regulatory or contractual deadline is irreversible on its date, a drifting model is causing harm every day it keeps running, and a portfolio review can absorb a provisional number. That ordering is defensible to all three requesters because it does not depend on your preferences.
  3. Take the deadline-bound work first, but scope it to the minimum defensible output, because completion factors feeding a filing carry a different error tolerance than a slide.
  4. Do not let the drifting model simply wait. Quantify the harm cheaply by comparing calibration in the affected region against the rest, and if it is materially miscalibrated propose flagging or suppressing its output for that region within the hour rather than at the end of a refresh.
  5. Give each deprioritised requester something real: a provisional forecast with its uncertainty and a refresh date, or a diagnostic that tells them whether their problem is urgent. Say no explicitly with a date rather than going quiet, because silence is what produces escalation.
Follow-up
  • The programme owner escalates to your manager. What do you want your manager to be able to say?
  • Midweek the actuarial work needs two more full days than you estimated. What gives?
  • How does your answer change if the drifting model drives a clinical outreach list rather than a report?
  • 01

    How do you approach disagreements with team members or stakeholders?

  • 02

    What motivates you to work in the healthcare sector?

  • 03

    Three requests land on Monday and you have one week. An actuarial team needs incurred-claims completion factors restated before a filing deadline on Thursday. A clinical programme owner wants a deterioration model refreshed because its calibration has drifted in one region. A trial operations team wants site enrolment forecasts for a portfolio review in two weeks. Each requester believes theirs is blocking. Deliverable: your sequence with the reasoning, the message you send to whoever is deprioritised, and the smaller artefact you hand each of the two you cannot fully serve.

PracHub interview preparation framework ↗
Is this an official Flatiron Health interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Flatiron Health. Rounds and questions reflect what candidates have reported, not a process Flatiron Health has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult are the interviews?

The interviews are designed to be challenging, reflecting the rigorous standards at Flatiron Health. Candidates should expect a mix of technical and behavioral questions that require thorough preparation.

PracHub interview research ↗
What differentiates successful candidates?

Successful candidates typically demonstrate a strong blend of technical expertise, problem-solving abilities, and a genuine passion for healthcare. They also exhibit effective communication skills and a collaborative mindset.

PracHub interview research ↗
What is the company culture like?

Flatiron Health fosters a culture of innovation, collaboration, and continuous learning. Employees are encouraged to share ideas and work together to drive improvements in healthcare.

PracHub interview research ↗
What is the typical timeline from application to offer?

The interview process can take several weeks, depending on the scheduling of interviews and assessments. Candidates should remain patient and proactive in their communication with the recruitment team.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.