University of Utah · Data Scientist
Updated · 2026-09-24

University of Utah Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

A Data Scientist at the University of Utah operates at the vital intersection of advanced quantitative analysis, clinical research, and academic innovation. Unlike typical tech-industry roles focused purely on commercial metrics, data science here directly impacts medical research, patient care outcomes, and public health initiatives. You will work closely with world-class faculty, clinical researchers, and biostatisticians to transform complex clinical and observational data into actionable scientific insights.

If the team owns experimentation, expect depth past a two-sample test: minimum detectable effect and its roughly inverse-square-root dependence on sample size (holding power, significance level and variance fixed), variance reduction from pre-period covariates, interference between units, and when a sequential design is the right call.

University of Utah candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Align cohorts to term weeks, not calendar weeksSeparate mastery signals from raw usage exposureReconstruct active time from raw heartbeat streams

31 min read

Practice 14 Data Scientist prompts
14Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Data Scientist at the University of Utah operates at the vital intersection of advanced quantitative analysis, clinical research, and academic innovation. Unlike typical tech-industry roles focused purely on commercial metrics, data science here directly impacts medical research, patient care outcomes, and public health initiatives. You will work closely with world-class faculty, clinical researchers, and biostatisticians to transform complex clinical and observational data into actionable scientific insights.

The role is highly collaborative and intellectually demanding. You will find yourself embedded within specialized departments, such as biostatistics units or academic medical centers, where your models and analyses will support clinical trials, grant proposals, and peer-reviewed publications. The insights you generate have a direct line of sight to improving human health, making the work both highly responsible and exceptionally rewarding.

To succeed in this position, you must possess a strong foundation in classical statistics, observational study design, and data management. You will navigate massive, sometimes messy, electronic health records and clinical databases. Your ability to write clean, reproducible code and explain complex statistical concepts to non-technical medical partners is what will set you apart as an invaluable member of the university's research community.

01

Recruiter Call

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Phone Interview

reported

A handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.

What to demonstrate

  • Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
  • Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
  • Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.

How to prepare

  • Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
  • Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
  • If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
PracHub interview research ↗
03

Timed Statistical Quiz

reported

Because the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.

What to demonstrate

  • Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
  • Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
  • Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs

How to prepare

  • For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
  • Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
  • For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
PracHub interview research ↗
04

Full-Day Interview

reported

An added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.

What to demonstrate

  • Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
  • Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
  • How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact

How to prepare

  • Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
  • For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
  • Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

The academic calendar creates structural breaks that look like product effects.

Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.

02

Consent and age rules silently truncate the data, not just the joins.

Where age_gated is TRUE, behavioural logging and cross-system joins are restricted, so those learners are missing from the very tables used to compute engagement. An analysis that simply inner-joins will produce a population skewed to older learners and self-serve accounts while appearing complete. Check the age_gated share of every population you report on, and state it.

03

Treating a non-significant result as proof of no effect

Say whether the confidence interval excludes the effect sizes you would have cared about. If it does not, the honest reading is that the test was underpowered, so report the minimum detectable effect the design could have found and what sample size would resolve it.

04

Reporting a p-value with no effect size or interval

Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

11 technical prompts3 include a worked solution

What are three distinct ways to reduce the width of a 95% confidence i…

medium
statistics and probability

What are three distinct ways to reduce the width of a 95% confidence interval in an observational study?

Approach
  1. Write down the assumption the method needs before you use the method.
  2. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  3. Sanity-check the answer against a simple bound or a simulated case.
Follow-up
  • Which assumption here is most likely to be violated in practice?
  • How would you explain this result to someone who does not know statistics?

How do you identify and handle confounding variables in a clinical dat…

medium
statistics and probability

How do you identify and handle confounding variables in a clinical data set?

Approach
  1. Say what the estimate is of, and over what population it generalises.
  2. Translate the result into the decision it informs, in one plain sentence.
  3. Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
  • What sample size would you need to detect an effect half this size?
  • Which assumption here is most likely to be violated in practice?

Imagine a hypothetical linear regression model with a dummy variable f…

medium
machine learning and modelling

Imagine a hypothetical linear regression model with a dummy variable for gender. How do you calculate and interpret the coefficient for male participants?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

Sessionise a raw heartbeat stream into activity rows

hardWorked solution
sessionisationevent streamsgroupbytime series

You get a DataFrame heartbeats with columns learner_id, content_item_id, content_version_no, and ts (UTC timestamps, unsorted). Produce one row per learner per session per content item with started_at, ended_at, wall_seconds, and active_seconds. A session closes after 30 minutes with no heartbeat from that learner. active_seconds sums consecutive gaps inside a session, capping each gap at 120 seconds. Attribute each gap to the content item of the earlier heartbeat. State your rule for a gap of exactly 1800 seconds and for a session with a single heartbeat.

Approach
  1. Sort by learner_id then ts, and compute the gap with a per-learner diff so the first heartbeat of each learner has no gap. A global diff lets one learner's last heartbeat set the next learner's first gap.
  2. Flag a session break where gap > 1800 seconds, then cumsum that boolean within the learner to get a session index. Say whether exactly 1800 opens a new session; either choice is defensible but it must be stated because it moves the session count.
  3. Each gap covers the interval before the current heartbeat, so attribute min(gap, 120) to the content_item_id of the previous row via shift, not the current row.
  4. Group by (learner_id, session_index, content_item_id): sum the attributed capped gaps for active_seconds, take min and max ts for started_at and ended_at, and derive wall_seconds from those two.
  5. Handle the degenerate group explicitly. A single heartbeat gives active_seconds 0 and wall_seconds 0. Dropping those rows quietly removes real opens and biases any completion or abandonment rate computed later.
  6. Validate that within every session the summed active_seconds never exceeds the session's wall_seconds, which catches attribution and cap errors in one assertion.
Worked solution 35 min
  1. Sort by (learner_id, ts) and compute gap = ts.diff().dt.total_seconds(), then null the gap on the first row of each learner using a learner-change mask.
  2. new_session = gap.isna() | (gap > 1800); session_index = new_session.groupby(learner_id).cumsum().
  3. contribution = gap.clip(upper=120).fillna(0), attributed to the previous row's content_item_id using shift within the learner and session.
  4. Aggregate by (learner_id, session_index, content_item_id) with sum for active_seconds and min and max of ts for the timestamps.
  5. Assert per session that sum(active_seconds) <= (max(ts) - min(ts)) in seconds, and print the count of single-heartbeat groups.
EXPECTED RESULTA frame keyed by (learner_id, session_index, content_item_id) where active_seconds equals the sum of min(gap, 120) over consecutive heartbeats with the first heartbeat contributing zero, wall_seconds equals ended_at minus started_at, and sum(active_seconds) <= wall_seconds holds for every session.
Follow-up
  • The client emits a heartbeat every 60 seconds on desktop and every 15 seconds on a managed device fleet. What does the 120-second cap do to the comparison between those two populations?
  • Your reconstructed active_seconds disagrees with the stored column by 4 percent overall. How do you localise the disagreement rather than argue about the total?
  • A learner has two tabs open on different items at once. What does your output claim, and what should it claim?

For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Build a fixture you can check answers against
  • Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
  • Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
  • Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.

Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Joins, filters and NULL semantics
  • Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
  • Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
  • Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.

Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.

Practice prompt ↗Practice prompt ↗
03Window functions and frames
  • Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
  • Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
  • Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.

Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.

Practice prompt ↗Practice prompt ↗
04The four analytical query patterns
  • Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
  • Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
  • Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.

Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Write SQL the way you will have to write it live
  • Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
  • Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
  • Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.

Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.

Practice prompt ↗Practice prompt ↗
06One day for everything that is not SQL
  • Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
  • Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
  • Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.

Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.

Practice prompt ↗Practice prompt ↗
07Full loop rehearsal
  • Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
  • Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
  • Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.

Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Work that nobody used is a common and unflattering pattern in data careers, and interviewers probe for it. Have a story about an analysis that changed a decision, and be specific about how you got it in front of the person who could act. Also have one about work that went nowhere, with your reading of why.

Describe a time when you had to work with a difficult stakeholder or f…

medium
behavioural and stakeholder questions

Describe a time when you had to work with a difficult stakeholder or faculty member. How did you align on the research methodology?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Name the disagreement or constraint, and how you resolved it with evidence.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What did you decide not to do, and why?
  • What would you do differently if you ran that project again?

Disagree with a product manager using evidence, not volume

medium
measurement validityadaptive systemsstakeholder disagreement

A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.

Approach
  1. Make the objection falsifiable before raising it. The claim implies two testable predictions: mean calibrated difficulty of served items rose with learner ability, and accuracy is flat within ability strata.
  2. Compute weekly mean irt_b of served items per arm, restricted to items with calibration_n above your floor, and plot it against the accuracy series. If served difficulty tracked ability, the accuracy line carries no learning signal and you can show that rather than assert it.
  3. Build the metric that survives adaptivity: a small fixed-form set with (content_item_id, version_no) held constant, served to both arms, reported as the pilot's accuracy readout.
  4. Bring the replacement to the meeting, not only the refutation. A PM who has been told the number is meaningless still has a launch decision and no instrument.
  5. Separate the two questions out loud: whether the selector helps learners is open and testable; whether first-attempt accuracy measures it is settled, and it does not.
Follow-up
  • The fixed-form set costs each learner six minutes a fortnight. How do you justify that to the same PM?
  • Mean served irt_b is flat but accuracy still rose four points. What do you look at next?

Account for an analysis you later discovered was wrong

hard
postmortemanalytical rigorownership

Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.

Approach
  1. Pick an error with a nameable mechanism, such as a grain mistake, a survivorship filter, or a join that silently dropped age_gated learners. 'I misread a chart' has nothing testable in it.
  2. State the blast radius in decisions rather than in dashboards: what was decided, what was spent or promised, and to whom.
  3. Describe detection honestly, including whether someone else caught it and how long it ran undetected. A slow external catch is a usable answer; a vague one is not.
  4. Give the disclosure order and the reason for it: the decision owner first, with the corrected number and its direction in the opening sentence, then the wider audience.
  5. Close on the systemic fix and say whether it has fired since. A control that has never triggered on anything is a resolution, not a fix, and saying so is worth more than claiming otherwise.
Follow-up
  • What did the corrected number change about the decision? If nothing changed, why did you escalate at all?
  • Has the check you added caught anything since, and on what?
  • What class of error would still get past you today?
  • 01

    Describe a time when you had to work with a difficult stakeholder or faculty member. How did you align on the research methodology?

  • 02

    A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.

  • 03

    Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.

PracHub interview preparation framework ↗
Is this an official University of Utah interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at University of Utah. Rounds and questions reflect what candidates have reported, not a process University of Utah has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How important is SAS compared to Python or R for this role?

While Python and R are highly valued for modern machine learning and data visualization, SAS remains a core legacy tool for many clinical and health outcomes databases at the University of Utah. Expect to be tested on basic SAS programming during the interview process, and be prepared to use it regularly on the job.

PracHub interview research ↗
What is the work environment and culture like?

The culture is collaborative, intellectual, and mission-driven. You will work alongside highly educated faculty and researchers who value scientific rigor over rapid commercial deployment. The pace is generally steadier than in the private tech sector, with a strong focus on work-life balance and professional development.

PracHub interview research ↗
How should I prepare for the full-day interview?

Practice explaining your past projects clearly, focusing on *why* you chose specific statistical methods. Be ready for a mix of technical exercises, behavioral questions, and casual conversations. Remember that the casual interactions, including lunch, are still part of the evaluation of your communication and collaboration skills.

PracHub interview research ↗
Is there flexibility for remote or hybrid work?

This varies by department and specific research group. Many data science and biostatistics teams at the university offer hybrid work arrangements, but some on-site presence is typically required to collaborate effectively with clinical faculty and attend departmental meetings.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.