A Data Scientist at the University of Utah operates at the vital intersection of advanced quantitative analysis, clinical research, and academic innovation. Unlike typical tech-industry roles focused purely on commercial metrics, data science here directly impacts medical research, patient care outcomes, and public health initiatives. You will work closely with world-class faculty, clinical researchers, and biostatisticians to transform complex clinical and observational data into actionable scientific insights.
The role is highly collaborative and intellectually demanding. You will find yourself embedded within specialized departments, such as biostatistics units or academic medical centers, where your models and analyses will support clinical trials, grant proposals, and peer-reviewed publications. The insights you generate have a direct line of sight to improving human health, making the work both highly responsible and exceptionally rewarding.
To succeed in this position, you must possess a strong foundation in classical statistics, observational study design, and data management. You will navigate massive, sometimes messy, electronic health records and clinical databases. Your ability to write clean, reproducible code and explain complex statistical concepts to non-technical medical partners is what will set you apart as an invaluable member of the university's research community.
Recruiter Call
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Phone Interview
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Timed Statistical Quiz
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Full-Day Interview
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
PracHub editorial advice for the preparation topics above.
The academic calendar creates structural breaks that look like product effects.
Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.
Consent and age rules silently truncate the data, not just the joins.
Where age_gated is TRUE, behavioural logging and cross-system joins are restricted, so those learners are missing from the very tables used to compute engagement. An analysis that simply inner-joins will produce a population skewed to older learners and self-serve accounts while appearing complete. Check the age_gated share of every population you report on, and state it.
Treating a non-significant result as proof of no effect
Say whether the confidence interval excludes the effect sizes you would have cared about. If it does not, the honest reading is that the test was underpowered, so report the minimum detectable effect the design could have found and what sample size would resolve it.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
What are three distinct ways to reduce the width of a 95% confidence i…
What are three distinct ways to reduce the width of a 95% confidence interval in an observational study?
Approach
- Write down the assumption the method needs before you use the method.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
How do you identify and handle confounding variables in a clinical dat…
How do you identify and handle confounding variables in a clinical data set?
Approach
- Say what the estimate is of, and over what population it generalises.
- Translate the result into the decision it informs, in one plain sentence.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
Imagine a hypothetical linear regression model with a dummy variable f…
Imagine a hypothetical linear regression model with a dummy variable for gender. How do you calculate and interpret the coefficient for male participants?
Approach
- Say how the offline result would be validated online before it is trusted.
- Check what information would not exist at prediction time, and exclude it.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Sessionise a raw heartbeat stream into activity rows
You get a DataFrame heartbeats with columns learner_id, content_item_id, content_version_no, and ts (UTC timestamps, unsorted). Produce one row per learner per session per content item with started_at, ended_at, wall_seconds, and active_seconds. A session closes after 30 minutes with no heartbeat from that learner. active_seconds sums consecutive gaps inside a session, capping each gap at 120 seconds. Attribute each gap to the content item of the earlier heartbeat. State your rule for a gap of exactly 1800 seconds and for a session with a single heartbeat.
Approach
- Sort by learner_id then ts, and compute the gap with a per-learner diff so the first heartbeat of each learner has no gap. A global diff lets one learner's last heartbeat set the next learner's first gap.
- Flag a session break where gap > 1800 seconds, then cumsum that boolean within the learner to get a session index. Say whether exactly 1800 opens a new session; either choice is defensible but it must be stated because it moves the session count.
- Each gap covers the interval before the current heartbeat, so attribute min(gap, 120) to the content_item_id of the previous row via shift, not the current row.
- Group by (learner_id, session_index, content_item_id): sum the attributed capped gaps for active_seconds, take min and max ts for started_at and ended_at, and derive wall_seconds from those two.
- Handle the degenerate group explicitly. A single heartbeat gives active_seconds 0 and wall_seconds 0. Dropping those rows quietly removes real opens and biases any completion or abandonment rate computed later.
- Validate that within every session the summed active_seconds never exceeds the session's wall_seconds, which catches attribution and cap errors in one assertion.
Worked solution 35 min
- Sort by (learner_id, ts) and compute gap = ts.diff().dt.total_seconds(), then null the gap on the first row of each learner using a learner-change mask.
- new_session = gap.isna() | (gap > 1800); session_index = new_session.groupby(learner_id).cumsum().
- contribution = gap.clip(upper=120).fillna(0), attributed to the previous row's content_item_id using shift within the learner and session.
- Aggregate by (learner_id, session_index, content_item_id) with sum for active_seconds and min and max of ts for the timestamps.
- Assert per session that sum(active_seconds) <= (max(ts) - min(ts)) in seconds, and print the count of single-heartbeat groups.
Follow-up
- The client emits a heartbeat every 60 seconds on desktop and every 15 seconds on a managed device fleet. What does the 120-second cap do to the comparison between those two populations?
- Your reconstructed active_seconds disagrees with the stored column by 4 percent overall. How do you localise the disagreement rather than argue about the total?
- A learner has two tabs open on different items at once. What does your output claim, and what should it claim?
Explain how you would structure a data pipeline to automate the cleani…
Explain how you would structure a data pipeline to automate the cleaning of incoming patient registry data.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Write a basic SAS program to merge two datasets and filter out missing…
Write a basic SAS program to merge two datasets and filter out missing observations.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
What is the difference between a merge and a join in SAS, and when wou…
What is the difference between a merge and a join in SAS, and when would you use less-than or greater-than operators to subset clinical trials?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Rebuild sessions and active seconds from heartbeat events
Given raw_heartbeat(learner_id, content_item_id, content_version_no, heartbeat_at TIMESTAMP UTC, session_hint VARCHAR NULL), reconstruct the grain of fct_lesson_activity: one row per learner per content item per session. Close a session after 30 minutes with no heartbeat. Compute active_seconds as the sum of inter-heartbeat gaps inside a session with each gap capped at 120 seconds, and wall_seconds as last minus first heartbeat. Ignore session_hint, which the client sets unreliably. Then reconcile your active_seconds against the stored column in fct_lesson_activity and explain any systematic difference.
Approach
- Compute prev_at = LAG(heartbeat_at) OVER (PARTITION BY learner_id, content_item_id, content_version_no ORDER BY heartbeat_at) and gap_seconds from the difference. Partition on the version too, since the same item at a new version is different content.
- Flag is_new_session = (prev_at IS NULL OR gap_seconds > 1800), then session_seq = SUM(is_new_session::int) OVER (same partition ORDER BY heartbeat_at ROWS UNBOUNDED PRECEDING). The running sum over an ordered flag is the island assignment and it stays correct at any data volume, unlike a self-join on nearby timestamps.
- Aggregate per (partition, session_seq): active_seconds = SUM(CASE WHEN is_new_session THEN 0 ELSE LEAST(gap_seconds, 120) END), wall_seconds = MAX(heartbeat_at) - MIN(heartbeat_at), plus heartbeat_count.
- State the two consequences of the definition before anyone asks: a single-heartbeat session has active_seconds = 0 and wall_seconds = 0, and the per-gap cap means active_seconds can never exceed 120 * (heartbeat_count - 1), so a client with a slow heartbeat interval understates real time.
- Reconcile by joining on (learner_id, content_item_id, started_at) and examining the distribution of the difference, not a row count. A uniform shift points at the cap or the close rule; a heavy one-sided tail points at heartbeats the stored column excludes.
Worked solution 30 min
- Write the LAG and gap CTE and eyeball ten rows for one busy learner before adding anything else.
- Add the is_new_session flag and the running-sum session_seq, then verify session_seq increments exactly where the gap exceeds 1800 seconds.
- Aggregate to the session grain with the CASE-guarded capped sum.
- Join to fct_lesson_activity on learner, item and started_at, and plot or bucket the signed difference in active_seconds.
Follow-up
- The stored column excludes backgrounded tabs but raw_heartbeat has no visibility flag. How would you detect backgrounded stretches from timing alone, and how confident should you be?
- Two heartbeats can share a timestamp. What does that do to LAG, and does it change the answer?
- A learner studies for two hours with a 40-minute dinner break in the middle. Your rule gives two sessions. Which downstream metrics care about that choice and which do not?
How do you prioritize your tasks when supporting multiple research pro…
How do you prioritize your tasks when supporting multiple research projects with competing grant deadlines?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Choose a success metric for autoplay lesson advance
A change autoplays the next lesson when a video ends. Week one: lessons opened per learner rose 33%, active minutes rose 8%, scored submissions per active learner fell from 2.1 to 2.0. You have fct_lesson_activity (active_seconds, wall_seconds, completion_status, is_assigned), dim_content_item (item_type, expected_minutes) and fct_assessment_response (attempt_no, is_correct, max_points). Name the primary metric you would have pre-registered, two guardrails, and state exactly how lessons opened and completion_status are moved by the mechanism rather than by learning.
Approach
- Sort every candidate metric into exposure or outcome. Autoplay decides what starts, so any metric whose numerator is a start, an open, or a video reaching its end is written by the feature itself and cannot be primary.
- Check whether active_seconds is also contaminated: it caps each inter-heartbeat gap at 120s and excludes backgrounded tabs, so a foreground autoplay with nobody watching still accrues time while a background one accrues none. Ask how the client emits heartbeats before quoting the 8%.
- Pick a primary metric that requires an action the feature does not perform for the learner: weekly learners with at least one scored submission, or verified mastery events per active learner where the delayed retention check is available for the grade band.
- Add guardrails that close the two obvious gaming routes: scored submissions per active learner reported alongside mean max_points per submission, and median active_seconds per completed (content_item_id, version_no) so click-through shows up as a drop.
- Decide against the pre-registered rule rather than against whichever number came back largest, and say that one week of data is a direction, not a result.
Worked solution 20 min
- List the four reported numbers and label each exposure or outcome, with one sentence per exposure metric describing how autoplay moves it at zero learning.
- Compute the outcome change: 2.0 versus 2.1 scored submissions per active learner is -4.8% relative, the only number in the set the feature does not write directly.
- Write the primary metric definition with its numerator, denominator and window, then the two guardrail definitions.
- Write the ship rule: ship only on a non-negative primary with both guardrails flat or better, holding the exposure metrics out of the decision entirely.
Follow-up
- What evidence would convince you the 8% active-minutes lift is attention rather than a foreground tab left open?
- Verified mastery is not calibrated for the youngest grade bands. What do you report there instead, and what do you lose?
- Autoplay may genuinely help learners who would have stopped at a natural break. How would you measure that group without slicing after seeing the results?
Net revenue retention slid nine points in one quarter
Net revenue retention on institutional licenses fell from 104 percent to 95 percent quarter over quarter. Source: fct_subscription_period (subscription_period_id, account_id, org_id, plan_tier, billing_interval, seats_purchased, seats_provisioned, mrr_usd, period_start, period_end, is_renewal, status, canceled_at, churn_reason). Decompose the nine points into expansion, contraction and non-renewal, and establish whether renewal behaviour changed at all. Deliverable: the decomposition, the cohort definition you used, and a call on whether this is a trend or account news.
Approach
- Pin the cohort in writing before computing anything: orgs holding an active period twelve months before each renewal date, new logos excluded from both numerator and denominator, and non-renewals entering the numerator as zero rather than being dropped. Dropping non-renewals is the most common way this metric is overstated, and correcting it can move the headline by more than the effect under investigation.
- Rebuild the ratio from components on that cohort: starting annualised revenue, expansion, contraction, non-renewal. Check the components sum back to the reported ratio in both quarters; a residual means the two quarters were not computed on the same cohort rule and the comparison is void before any interpretation.
- Test cohort comparability on billing_interval. Multi-year contracts do not present a renewal in every window, so a quarter whose renewal cohort holds a different multi_year share is not comparable to its predecessor. Recompute both quarters holding the interval mix fixed and report that alongside the raw number.
- Rank per-org contributions to the nine points and state how much the top three carry. Institutional revenue is concentrated, so a decomposition that leaves the concentration unstated invites a trend reading of what is account news.
- Audit the inputs: confirm mrr_usd normalisation is consistent (annual divided by twelve) and look for overlapping period_start and period_end within a single account_id, which double counts revenue after a mid-term upgrade.
- Close the loop to something actionable by joining the lost orgs back to seat activation and pace adherence in the preceding term, and say whether the signal was visible early enough for anyone to have acted on it.
Follow-up
- Build the renewal-risk view from this: which signals are actionable while there is still time to act, and at what days-to-renewal would you surface an org?
- An org contracts from 900 seats to 600 but moves to a higher tier and raises MRR. Expansion or contraction, and does your decomposition handle it without double counting?
- How would you present NRR so that a quarter with an unusual renewal cohort is not read as a trend by an executive audience?
For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Build a fixture you can check answers against
- Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
- Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
- Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.
Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Joins, filters and NULL semantics
- Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
- Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
- Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.
Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.
Practice prompt ↗Practice prompt ↗03Window functions and frames
- Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
- Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
- Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.
Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.
Practice prompt ↗Practice prompt ↗04The four analytical query patterns
- Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
- Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
- Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.
Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Write SQL the way you will have to write it live
- Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
- Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
- Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.
Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.
Practice prompt ↗Practice prompt ↗06One day for everything that is not SQL
- Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
- Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
- Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.
Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.
Practice prompt ↗Practice prompt ↗07Full loop rehearsal
- Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
- Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
- Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.
Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Work that nobody used is a common and unflattering pattern in data careers, and interviewers probe for it. Have a story about an analysis that changed a decision, and be specific about how you got it in front of the person who could act. Also have one about work that went nowhere, with your reading of why.
Describe a time when you had to work with a difficult stakeholder or f…
Describe a time when you had to work with a difficult stakeholder or faculty member. How did you align on the research methodology?
Approach
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that project again?
Disagree with a product manager using evidence, not volume
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
Approach
- Make the objection falsifiable before raising it. The claim implies two testable predictions: mean calibrated difficulty of served items rose with learner ability, and accuracy is flat within ability strata.
- Compute weekly mean irt_b of served items per arm, restricted to items with calibration_n above your floor, and plot it against the accuracy series. If served difficulty tracked ability, the accuracy line carries no learning signal and you can show that rather than assert it.
- Build the metric that survives adaptivity: a small fixed-form set with (content_item_id, version_no) held constant, served to both arms, reported as the pilot's accuracy readout.
- Bring the replacement to the meeting, not only the refutation. A PM who has been told the number is meaningless still has a launch decision and no instrument.
- Separate the two questions out loud: whether the selector helps learners is open and testable; whether first-attempt accuracy measures it is settled, and it does not.
Follow-up
- The fixed-form set costs each learner six minutes a fortnight. How do you justify that to the same PM?
- Mean served irt_b is flat but accuracy still rose four points. What do you look at next?
Account for an analysis you later discovered was wrong
Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.
Approach
- Pick an error with a nameable mechanism, such as a grain mistake, a survivorship filter, or a join that silently dropped age_gated learners. 'I misread a chart' has nothing testable in it.
- State the blast radius in decisions rather than in dashboards: what was decided, what was spent or promised, and to whom.
- Describe detection honestly, including whether someone else caught it and how long it ran undetected. A slow external catch is a usable answer; a vague one is not.
- Give the disclosure order and the reason for it: the decision owner first, with the corrected number and its direction in the opening sentence, then the wider audience.
- Close on the systemic fix and say whether it has fired since. A control that has never triggered on anything is a resolution, not a fix, and saying so is worth more than claiming otherwise.
Follow-up
- What did the corrected number change about the decision? If nothing changed, why did you escalate at all?
- Has the check you added caught anything since, and on what?
- What class of error would still get past you today?
- 01
Describe a time when you had to work with a difficult stakeholder or faculty member. How did you align on the research methodology?
- 02
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
- 03
Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.
Is this an official University of Utah interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at University of Utah. Rounds and questions reflect what candidates have reported, not a process University of Utah has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How important is SAS compared to Python or R for this role?
While Python and R are highly valued for modern machine learning and data visualization, SAS remains a core legacy tool for many clinical and health outcomes databases at the University of Utah. Expect to be tested on basic SAS programming during the interview process, and be prepared to use it regularly on the job.
PracHub interview research ↗What is the work environment and culture like?
The culture is collaborative, intellectual, and mission-driven. You will work alongside highly educated faculty and researchers who value scientific rigor over rapid commercial deployment. The pace is generally steadier than in the private tech sector, with a strong focus on work-life balance and professional development.
PracHub interview research ↗How should I prepare for the full-day interview?
Practice explaining your past projects clearly, focusing on *why* you chose specific statistical methods. Be ready for a mix of technical exercises, behavioral questions, and casual conversations. Remember that the casual interactions, including lunch, are still part of the evaluation of your communication and collaboration skills.
PracHub interview research ↗Is there flexibility for remote or hybrid work?
This varies by department and specific research group. Many data science and biostatistics teams at the university offer hybrid work arrangements, but some on-site presence is typically required to collaborate effectively with clinical faculty and attend departmental meetings.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22