As a Data Scientist at Headway, you are not just a builder of models or dashboards; you are a core architect of the truth engine that powers the future of mental healthcare. Headway is scaling rapidly, moving beyond its roots in automated insurance billing to become the primary platform where therapy occurs. In this role, you will define how the organization measures success across product surfaces, growth channels, and clinical outcomes.
You will face high-stakes, ambiguous problems where the signal is often noisy and the stakes involve real patient access to care. Whether you are designing Bayesian experimentation frameworks, diagnosing sudden shifts in product metrics, or architecting causal inference models, your work directly informs strategy, policy, and resource allocation. At Headway, Data Scientists are expected to zoom in to debug data pipelines and zoom out to influence company-wide product strategy, making this a high-impact, leadership-oriented position.
Recruiter Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Rounds
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Cross-Functional Interviews
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Final Leadership Discussions
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
7 candidate reports. Individual accounts describe a particular role and hiring cycle.
Headway Software Engineer interview: withdrew after a delayed process
I initially had an easier path into the process, but it became frustrating quickly. The recruiter seemed unprepared, and the timeline took too long to get moving. Eventually, I decided it wasn’t worth my time and withdrew to focus on other interviews. The process felt unprofessional to me. I didn’t like how long things dragged before anything substantial happened. Because I withdrew and didn’t co…
Read full experienceHeadway Customer Success Engineer interview with an abrupt manager round
I was contacted through LinkedIn, and someone from the recruiting team arranged a recruiter screen. That interview was fairly normal, covering my work experience along with an overview of the company and role. Afterward, I moved straight to a hiring manager interview that felt very different. The hiring manager barely introduced themselves, didn't ask much about me, and instead went straight to a…
Read full experienceHeadway Software Engineer interview with role homework and employee interviews
I went through a multi-step process that felt unusual in how quickly it ramped up. It started with a recruiter meeting where I answered basic questions. Then I was assigned homework tied to the role, which I had to complete before the next technical conversation. After that, I had a technical interview with the hiring manager. The format shifted again when I interviewed with current employees. Th…
Read full experienceHeadway Software Engineer interview: unclear onsite design expectations
My process started with a recruiter touchpoint, followed by a technical screen and then a three-round onsite. Scheduling and feedback were responsive, but the actual evaluation didn't line up with what I expected. The onsite design segment especially didn't feel like a classic system design round. The constraints made it hard to tell what direction they wanted, so I spent more time guessing than…
Read full experienceHeadway Backend Engineer interview: recruiter screen and unclear follow-up
My process started with a quick recruiter screen focused on behavioral questions. We covered the usual background topics, along with how I’d troubleshoot technical issues when something breaks and how I’d explain a technical concept to someone without a technical background. I made it into the hiring process, but it ended without an offer. In one experience, the recruiter interaction had a good t…
Read full experiencePracHub editorial advice for the preparation topics above.
Pre-post evaluation on a high-cost or high-risk cohort
Cohorts selected on an extreme value of the outcome regress toward the mean on their own. Members identified as the top 1 percent of spend in one year spend far less the next, whether or not anyone intervenes, because the selecting year captured both chronic severity and one-off events. A pre-post design on such a cohort will report savings every time. A concurrent comparison group selected by the same rule in the same period, or a regression discontinuity at the selection threshold, is the minimum credible design.
Immortal time in adherence, treatment, and enrolment definitions
Classifying members as adherent, treated, or programme-enrolled requires them to survive and stay covered long enough to accumulate the defining events. That guaranteed event-free interval is assigned to the exposed group, so the exposure looks protective for reasons that have nothing to do with the treatment. Adherence studies are the classic case: measuring 12-month proportion of days covered and then comparing mortality builds survival into the exposure definition. Use time-varying exposure or a landmark analysis with the classification window excluded from follow-up.
Naming a model class before naming the deployment constraints
Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.
Reaching for a model before the target metric exists
Before naming an algorithm, write down the label, the prediction time, and the action that changes when the score crosses a threshold. If you cannot say what decision the output drives, any modelling choice is guesswork dressed up as method.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
If we launch a new subscription model for therapists, how do we measur…
If we launch a new subscription model for therapists, how do we measure the cannibalization of existing services?
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Simulate how ignoring clustering inflates the false positive rate
Simulate a provider-level comparison with no true effect. 40 providers, 60 members each, a continuous outcome with total variance 1 and intraclass correlation 0.02. Randomise providers 20 to 20, not members. Over 5,000 replications, report the share of replications where a two-sample t-test run on all 2,400 individual observations, ignoring provider, returns p below 0.05. Report the same share when the unit of analysis is the 40 provider means. Then state how the first number follows from the design effect 1 + (m - 1) x ICC.
Approach
- Decompose the variance explicitly rather than tuning it: provider variance = ICC = 0.02, residual variance = 1 - ICC = 0.98. Draw u_j once per provider and e_ij per member, so the outcome is u_j + e_ij and the correlation between two members of one provider is 0.02 by construction.
- Randomise at the provider level. That is the entire mechanism: the treatment indicator is constant within a provider, so in any single replicate the provider intercepts are confounded with the arm, and the naive test reads that confounding as signal.
- Vectorise over replications with a (reps, providers, members) array. 5,000 x 2,400 normals is trivial in numpy, and a Python loop is what tempts people to cut replications to 500 and report a number with Monte Carlo noise of plus or minus 0.017.
- Predict the answer before running it: design effect = 1 + (60 - 1) x 0.02 = 2.18, naive standard errors are too small by sqrt(2.18) = 1.476, so the rejection rate at nominal 0.05 is about 2 x (1 - Phi(1.96 / 1.476)) = 0.18. Agreement between prediction and simulation is what makes the result a demonstration rather than an anecdote.
- Run the provider-mean analysis as the control. Recovering 0.05 there proves the generator is correct and isolates the failure to the analysis unit.
Worked solution 30 min
- Draw u with shape (reps, 40) from N(0, sqrt(0.02)) and e with shape (reps, 40, 60) from N(0, sqrt(0.98)); y = u[:, :, None] + e.
- Assign providers 0-19 to one arm and 20-39 to the other; because there is no true effect, a fixed split is valid and removes one source of Monte Carlo noise.
- Naive test: flatten each arm to 1,200 observations and compute a two-sample t statistic per replication with vectorised means and pooled variance.
- Cluster test: average within provider to 40 values, then run the same two-sample t on 20 versus 20.
- Report both rejection shares and compare the naive one against 2 x (1 - Phi(1.96 / sqrt(2.18))).
Follow-up
- How many providers would you need for 80 percent power on a 0.2 SD difference with 60 members each and this ICC?
- Cluster sizes are unequal in reality. The design effect approximation becomes 1 + ((1 + CV^2) x mbar - 1) x ICC. Which direction does that move your sample size and why?
- The intervention cannot be withheld from any provider. Sketch a design that still yields a defensible estimate.
Proportion of days covered with shifted, truncated refill intervals
pharmacy_claim has fill_id, member_id, therapeutic_class_code, fill_date, days_supply, reversal_flag, reversed_fill_id. For one therapeutic class and a fixed 12-month window, compute proportion of days covered per member: distinct days on which a dispensed days_supply covers the day, divided by days from the member's first in-window fill through the window end. An early refill shifts coverage forward rather than stacking, and coverage is truncated at the window end. Drop both rows of every reversed pair. Return member_id, pdc, and the share at pdc >= 0.80 among members with at least 2 fills and at least 91 days of follow-up.
Approach
- Remove reversals as pairs first. Drop every row with reversal_flag true, and also drop the fill_ids those rows point at through reversed_fill_id. Dropping only the flagged row leaves a dispense that was never collected in the exposure.
- Walk fills per member in fill_date order carrying a cursor: start = max(fill_date, previous_end + 1 day), end = start + days_supply - 1. The shift is path dependent, so a plain cumsum over days_supply does not reproduce it. Use itertools.accumulate or a per-member loop over numpy arrays, not a row-wise apply over the whole frame.
- Truncate the last interval at the window end before measuring. Without truncation a 90-day fill dispensed on the final day pushes the covered-day count past the denominator and PDC above 1.0.
- Use the denominator the definition states: first in-window fill_date through window end, inclusive. Not a flat 365, and not first fill to last fill, which is a different metric that rewards early discontinuation.
- Apply the eligibility filter before computing the >= 0.80 share, and report the size of that denominator next to the share. A share without its denominator is not reviewable.
Follow-up
- How does PDC differ from medication possession ratio, and which of the two can exceed 1.0?
- A member switches to a different ingredient inside the same therapeutic class mid-window. Should the intervals chain, and what does that do to the class-level number?
- Members who die or lose coverage mid-window get a short denominator and often a high PDC. If you then compare mortality by adherence category, what bias have you built in and how do you remove it?
Write a SQL query using window functions to calculate the rolling 30-d…
Write a SQL query using window functions to calculate the rolling 30-day retention rate for providers.
Approach
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Given a table of daily active usage, write a query to find the top 3 c…
Given a table of daily active usage, write a query to find the top 3 cohorts by growth rate.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Explain how you would optimize a query that is joining a massive claim…
Explain how you would optimize a query that is joining a massive claims table with a sparse patient activity table.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Find inpatient encounters with no resulted lab during the stay
encounter holds encounter_id, patient_id, encounter_type, admit_ts, discharge_ts and facility_id. lab_result holds result_id, order_id, patient_id, encounter_id (NULL for outpatient standing orders), loinc_code, specimen_collected_ts, resulted_ts and result_status. Return every inpatient encounter discharged in 2025 with no final or corrected lab_result linked to it. A colleague wrote WHERE encounter_id NOT IN (SELECT encounter_id FROM lab_result) and got zero rows back. Say why, write the correct query, then say what changes when labs must instead be matched on patient_id with the specimen collected between admit_ts and discharge_ts.
Approach
- Name the mechanism rather than the symptom. x NOT IN (subquery) expands to x <> y1 AND x <> y2 AND ..., and a single NULL y makes that chain UNKNOWN, never TRUE. Since lab_result.encounter_id is nullable by design for standing orders, this query can only ever return zero rows.
- Rewrite as NOT EXISTS, or LEFT JOIN with an IS NULL filter. Both treat a NULL key as no match rather than as unknown. Adding IS NOT NULL inside the NOT IN subquery also works, but NOT EXISTS is the habit that survives someone making another column nullable later.
- Put the result_status predicate inside the correlated subquery or the ON clause, not in an outer WHERE. Outside, it turns the anti-join back into an inner join and an encounter whose only labs were cancelled disappears instead of qualifying.
- For the timestamp variant, match on patient_id with specimen_collected_ts inside the stay rather than resulted_ts. A specimen drawn an hour before discharge can result the next day, and anchoring on resulted_ts would wrongly call that stay lab-free.
- Restrict to encounter_type 'inpatient' and discharge_ts in 2025, and report open encounters with a NULL discharge_ts as a separate count instead of letting the filter swallow them.
Worked solution 20 min
- Demonstrate the cause: SELECT COUNT(*) FROM lab_result WHERE encounter_id IS NULL returns a non-zero number.
- Write the NOT EXISTS version with result_status IN ('final','corrected') inside the correlated subquery.
- Write the LEFT JOIN version with the same predicate in the ON clause and WHERE l.result_id IS NULL, then compare counts.
- Build the timestamp variant joining on patient_id with specimen_collected_ts >= admit_ts AND < discharge_ts.
- Report both counts side by side and explain which encounters differ between the two definitions.
Follow-up
- Write the LEFT JOIN form and say exactly where the result_status predicate must sit for the two forms to agree.
- How would you separate encounters whose only labs were cancelled from encounters with no lab rows at all?
- Some encounters carry member_id NULL because the person index did not match. What does that do to a payer-side version of this measure?
How do you balance long-term clinical outcomes with short-term platfor…
How do you balance long-term clinical outcomes with short-term platform engagement metrics?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
We noticed a 5% drop in session bookings last week; describe your step…
We noticed a 5% drop in session bookings last week; describe your step-by-step process for diagnosing the root cause.
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How would you define the "North Star" metric for a new patient-provide…
How would you define the "North Star" metric for a new patient-provider matching feature?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
What are the most common experimentation pitfalls you have encountered…
What are the most common experimentation pitfalls you have encountered, and how do you mitigate them?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
How do you decide between a frequentist and a Bayesian approach for a …
How do you decide between a frequentist and a Bayesian approach for a long-running product experiment?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
Power a readmission trial on a rare binary outcome
A care-transitions programme (pharmacist call plus a seven-day follow-up visit) will be tested on index inpatient discharges drawn from encounter, excluding planned admissions, acute-to-acute transfers, AMA discharges, and dispositions of expired or hospice. Baseline 30-day unplanned readmission on eligible stays is 15.8 percent. Leadership wants to detect a 1.5 percentage point absolute reduction, two-sided alpha 0.05, 80 percent power, 1:1 individual randomisation at discharge. Give the required index stays per arm, and the minimum detectable absolute effect if only 4,000 eligible stays per arm accrue in the study window.
Approach
- Write the two-proportion sample size formula and state its preconditions: normal approximation, two-sided test, equal allocation, and the outcome measured on index stays rather than on members, since one member can contribute several index stays.
- Plug in p1 = 0.158, p2 = 0.143, and (z_0.975 + z_0.80)^2 = (1.960 + 0.842)^2 = 7.849. Report the answer as index stays, and separately as the calendar time it implies given the observed monthly eligible volume.
- Invert the same formula for the fixed-N case: solve for delta with n = 4,000 per arm, and say whether you used pooled or arm-specific variance because the two answers differ by about 0.1 percentage points.
- Translate the MDE into the relative reduction it implies and hand it back to leadership before the study starts, so the question becomes whether that effect is worth running rather than whether the null result was a failure.
- Flag the two adjustments that will make the real number larger: if a member can appear more than once, cluster on member_id, and if the programme is delivered by unit or by discharging team rather than to individuals, an individual-level calculation is the wrong model entirely.
Worked solution 20 min
- Compute (1.960 + 0.842)^2 = 7.849.
- Compute the variance sum: 0.158 x 0.842 = 0.1330, 0.143 x 0.857 = 0.1226, total 0.2556.
- n per arm = 7.849 x 0.2556 / 0.015^2 = 7.849 x 1,136 = 8,916, so round to about 8,920 per arm.
- For the MDE at n = 4,000, solve delta = sqrt(7.849 x variance sum / 4,000); with arm-specific variance this gives 0.0221, with pooled variance at p = 0.158 it gives 0.0229.
- Express the MDE as a relative effect: 0.023 / 0.158 is about a 15 percent relative reduction.
Follow-up
- The same programme is delivered by discharge unit, not to individual patients. What changes in the calculation?
- How would you handle a member who has three eligible index stays during the study window?
- Leadership proposes powering on a composite of readmission or ED revisit instead. What does that buy and what does it cost?
Allowed PMPM fell fourteen percent in three months
Allowed PMPM in the monthly series fell from $412 to $354 over the three most recent incurred months, and leadership wants to announce the saving. You have medical_claim_line (allowed_amount, service_start_date, received_date, paid_date, claim_id, claim_version, frequency_code, place_of_service_code), pharmacy_claim (allowed_amount, fill_date, reversal_flag) and member_enrollment coverage spans. The chart carries no paid-through date. Work out how much of the fall survives completion, and hand back a restated series with an explicit paid-through date and the completion factors you applied.
Approach
- Get the paid-through date first and put it on the chart. Without it the series is uninterpretable, because every incurred month is a different age.
- Build a lag triangle: for each incurred month, cumulative allowed_amount by the number of months between service_start_date and paid_date, on surviving claim versions only (drop frequency_code 8 voids and keep the highest claim_version per claim_id).
- Estimate development factors separately for pharmacy fills, professional lines and inpatient facility lines. Pharmacy adjudicates within days, professional within weeks, inpatient facility over months, so one blended factor understates the correction on the newest month and overstates it on the oldest.
- Gross each of the last three incurred months up by the reciprocal of its cumulative development factor, using factors fitted on months that have already fully run out.
- Rebuild the denominator independently as member-months from coverage spans. Member-months are complete on day one, so if the denominator also shows a lag pattern the enrollment file is late, which is a different bug with a different fix.
- Restate the series, mark the last three months as estimated, and report the residual movement after completion as the only part worth investigating.
Follow-up
- The completed series still shows a three percent fall. What do you look at next, and in what order?
- How would you detect that the fall is a mix shift toward cheaper services rather than lower volume?
- A contract settles on this number at a fixed paid-through date. What do you owe the other party about the estimate you just made?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Nearly every data role forces a trade between the analysis you want and the one that fits the decision window. Prepare a case where you deliberately shipped something less rigorous, named the weakness to the person relying on it, and said what would change your answer. The naming is the part interviewers listen for.
Describe a time you had to pivot a product roadmap based on a counter-…
Describe a time you had to pivot a product roadmap based on a counter-intuitive data finding.
Approach
- Pick a story where you drove the decision, not one where you observed it.
- State the situation in two sentences and spend the rest on your reasoning.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know the outcome was caused by your change?
Scope a one-line request for a readmission rate
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
Approach
- Read the probe: this tests whether you convert ambiguity into a specification without either stalling or guessing silently. Both failure modes are common, and the silent guess is worse because nobody can see it.
- Rank your questions by leverage on the number rather than by curiosity. Whether the numerator is all-cause or unplanned, and whether the clock starts at discharge_ts, move the result far more than a tie-break rule on same-day returns.
- Ask about the comparison first. A number going to a committee will be compared to something, and whether that something is last quarter, another panel, or a national figure decides whether you owe them a raw rate, a risk-adjusted rate, or an observed-over-expected ratio.
- Commit to a default in writing so silence does not block you: unplanned acute inpatient readmission within 30 days of discharge_ts, index stays excluding planned admissions, acute-to-acute transfers, discharges against medical advice, and dispositions of expired or hospice, denominator of eligible index discharges, paid-through date stated.
- Write the caveat as a restriction on use rather than a hedge: name the one comparison the figure supports and the one it does not.
Follow-up
- They reply that they want it 'the way the benchmark does it'. What do you ask next?
- The committee wants it split by attending provider. What changes in the specification, and what do you refuse to show?
- How does your answer change if the meeting is tomorrow rather than Friday?
Own an analysis that shipped wrong and was acted on
Describe an analysis you delivered that was wrong, where someone acted on it before the error surfaced. Pick something with a mechanism you can name, not a typo. Cover what the number claimed, what was actually true, how much time passed before it was caught, who caught it, and what you told the people who had already acted. Deliverable: a five-minute account that ends with the specific control you added afterwards, plus one occasion since when that control has fired.
Approach
- Read the probe: the interviewer is testing whether you detect your own errors and whether the correction was structural. Pick an error with a nameable mechanism (a denominator built from distinct members instead of member-months, a series read before claims runout completed, a cohort selected on the outcome) so the story has a diagnosis rather than a mood.
- Open with the decision that was made on your number, not with the bug. The severity of an analytic error is measured in the action it caused: a rate filing, a staffing plan, a programme expansion.
- State the counterfactual number. 'The true figure was 14.1 admissions per 1,000 member-years, not 9.8' is a fact; 'it was off by a lot' is not.
- Describe the detection path honestly, including whether you found it or somebody else did. Say how long it took, because time to detection is the part you actually control next time.
- Close on the control, and make it something that protects a reader who has never met you: a unit test on the denominator, a paid-through date stamped on every chart, a required exclusions block in the spec template. Name one later date it caught something.
Follow-up
- What would have had to be true for you to catch this before delivery rather than after?
- Did the stakeholder change how much they trust your work, and did you want them to?
- Give an example of a control you added that turned out to be the wrong control.
- 01
Describe a time you had to pivot a product roadmap based on a counter-intuitive data finding.
- 02
A clinical operations lead messages you: 'What is our readmission rate? The committee meets Friday.' You have encounter (encounter_id, admit_ts, discharge_ts, encounter_type, admission_type, discharge_disposition, drg_code, index_encounter_id) and member_enrollment coverage spans. Do not open a query editor yet. Deliverable: the three to five questions you send back, ranked by how much the answer moves the number; the default specification you will build if nobody replies by Wednesday; and the one line you will place under the figure so the committee does not read it against a published benchmark.
- 03
Describe an analysis you delivered that was wrong, where someone acted on it before the error surfaced. Pick something with a mechanism you can name, not a typo. Cover what the number claimed, what was actually true, how much time passed before it was caught, who caught it, and what you told the people who had already acted. Deliverable: a five-minute account that ends with the specific control you added afterwards, plus one occasion since when that control has fired.
Is this an official Headway interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Headway. Rounds and questions reflect what candidates have reported, not a process Headway has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews?
The technical rounds are challenging but practical. You will not be asked for obscure algorithms; instead, focus on your ability to write clean, efficient SQL and explain the "why" behind your analytical choices.
PracHub interview research ↗Does Headway focus more on product or machine learning?
This role is heavily biased toward Product Analytics and Experimentation. While Machine Learning is a component, the primary need is for rigorous decision support, causal inference, and measurement.
PracHub interview research ↗What is the culture like for Data Scientists?
Headway is a mission-driven company. The work is high-stakes, and you will be expected to move quickly. The environment is collaborative but demands high individual ownership over your analytical outputs.
PracHub interview research ↗How can I stand out?
Successful candidates demonstrate a "product-first" mindset. When answering any technical question, always connect your answer back to the business impact and the user experience.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22