As a Data Scientist at the University of Michigan, you will play a pivotal role in harnessing data to drive insights that influence critical decisions across various departments and initiatives. This position is vital to advancing the university's mission of education, research, and community engagement by utilizing data analytics to improve student outcomes, enhance operational efficiency, and support groundbreaking research initiatives.
You will be involved in projects that span multiple domains, such as education analytics, healthcare research, and urban planning, providing you with opportunities to make a significant impact on both local and global scales. The complexity and scale of the data you will work with are substantial, often requiring innovative modeling, machine learning techniques, and advanced statistical analyses. This role not only demands technical expertise but also the ability to communicate findings effectively to diverse stakeholders, making it a stimulating and rewarding position.
Prescreening Telephonic Interview
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
Online Interview
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Technical Interview
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Discussion with HR
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
PracHub editorial advice for the preparation topics above.
The academic calendar creates structural breaks that look like product effects.
Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.
Adaptive item selection holds observed accuracy flat by construction.
A selector targeting a fixed success probability, say 70%, will keep measured first-attempt accuracy near 70% whether learners are improving or not, because as theta rises the engine simply serves harder items. Reporting 'accuracy improved 3 points' on adaptive content therefore usually means the selector got more conservative, not that anyone learned more. Accuracy is only interpretable on a fixed form with unchanged item versions, which is why the first-attempt accuracy metric above carries both restrictions.
Analysing at a different unit than the one randomised
Say out loud what was randomised (user, device, account, cluster) and make the analysis unit match, or account for the clustering with cluster-robust standard errors, the delta method, or aggregation up to the randomised unit. Randomising users and then running a test over sessions understates variance and inflates the false-positive rate.
Explaining an aggregate move without decomposing the mix shift
Split the change in the aggregate into within-segment movement and movement in segment weights before you explain it. Every segment's rate can fall while the overall rate rises, purely because volume shifted toward segments that already had higher rates.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Write a function to calculate the mean and standard deviation of a lis…
Write a function to calculate the mean and standard deviation of a list of numbers.
Approach
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
What is the significance of p-values in hypothesis testing?
What is the significance of p-values in hypothesis testing?
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Translate the result into the decision it informs, in one plain sentence.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
Discuss a machine learning project you worked on and the challenges yo…
Discuss a machine learning project you worked on and the challenges you faced.
Approach
- Check what information would not exist at prediction time, and exclude it.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- What would you monitor after launch to know the model is still valid?
- How would you choose the decision threshold, and who owns that choice?
Attach the calibration in force without merge_asof
You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.
Approach
- Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
- Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
- Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
- Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
- Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
- Take the parameters positionally with np.take and write NaN where any invalidation fired.
Worked solution 30 min
- codes = pd.factorize on the concatenated item ids, applied to both frames so the mapping is shared.
- cal_key = cal_code.astype('int64') * (1 << 32) + valid_from_seconds; resp_key built the same way; sort cal by cal_key.
- pos = np.searchsorted(cal_key, resp_key, side='right') - 1.
- valid = (pos >= 0) & (cal_code[pos] == resp_code) & (cal_valid_to_seconds[pos].isna() | (resp_seconds < cal_valid_to_seconds[pos])).
- Assign irt_a and irt_b via np.take(pos) where valid, NaN elsewhere, and report the NaN count by reason.
Follow-up
- A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
- The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
- How would you test this without a golden output to compare against?
Rebuild sessions and active seconds from heartbeat events
Given raw_heartbeat(learner_id, content_item_id, content_version_no, heartbeat_at TIMESTAMP UTC, session_hint VARCHAR NULL), reconstruct the grain of fct_lesson_activity: one row per learner per content item per session. Close a session after 30 minutes with no heartbeat. Compute active_seconds as the sum of inter-heartbeat gaps inside a session with each gap capped at 120 seconds, and wall_seconds as last minus first heartbeat. Ignore session_hint, which the client sets unreliably. Then reconcile your active_seconds against the stored column in fct_lesson_activity and explain any systematic difference.
Approach
- Compute prev_at = LAG(heartbeat_at) OVER (PARTITION BY learner_id, content_item_id, content_version_no ORDER BY heartbeat_at) and gap_seconds from the difference. Partition on the version too, since the same item at a new version is different content.
- Flag is_new_session = (prev_at IS NULL OR gap_seconds > 1800), then session_seq = SUM(is_new_session::int) OVER (same partition ORDER BY heartbeat_at ROWS UNBOUNDED PRECEDING). The running sum over an ordered flag is the island assignment and it stays correct at any data volume, unlike a self-join on nearby timestamps.
- Aggregate per (partition, session_seq): active_seconds = SUM(CASE WHEN is_new_session THEN 0 ELSE LEAST(gap_seconds, 120) END), wall_seconds = MAX(heartbeat_at) - MIN(heartbeat_at), plus heartbeat_count.
- State the two consequences of the definition before anyone asks: a single-heartbeat session has active_seconds = 0 and wall_seconds = 0, and the per-gap cap means active_seconds can never exceed 120 * (heartbeat_count - 1), so a client with a slow heartbeat interval understates real time.
- Reconcile by joining on (learner_id, content_item_id, started_at) and examining the distribution of the difference, not a row count. A uniform shift points at the cap or the close rule; a heavy one-sided tail points at heartbeats the stored column excludes.
Follow-up
- The stored column excludes backgrounded tabs but raw_heartbeat has no visibility flag. How would you detect backgrounded stretches from timing alone, and how confident should you be?
- Two heartbeats can share a timestamp. What does that do to LAG, and does it change the answer?
- A learner studies for two hours with a 40-minute dinner break in the middle. Your rule gives two sessions. Which downstream metrics care about that choice and which do not?
Consecutive instructional-week streaks and term-over-term return
Using fct_lesson_activity(learner_id, enrollment_id, activity_date, active_seconds) and a calendar dim_term_week(term_id, term_week_no, week_start_date, week_end_date, is_instructional), where holiday weeks carry is_instructional = FALSE, compute for each learner and term the longest run of consecutive instructional weeks containing at least one activity day. A non-instructional week must not break a run, and activity inside one neither extends nor starts a run. Then compute term-over-term return rate into term T+1 with a floor of five activity days in term T. Return learner_id, term_id, activity_days counted over the whole term including non-instructional weeks, longest_streak_weeks and the return flag.
Approach
- Collapse activity to (learner_id, term_id, term_week_no) with COUNT(DISTINCT activity_date). fct_lesson_activity is one row per item per session, so a learner with twelve rows on one afternoon is one activity day and counting rows inflates everything downstream.
- Re-index before differencing. Apply DENSE_RANK() OVER (PARTITION BY term_id ORDER BY term_week_no) across instructional weeks only, producing a gap-free sequence where a holiday week has simply been removed. Differencing raw term_week_no instead splits a run at every holiday, which is exactly the artefact the task forbids.
- Accept what dropping those weeks costs. Activity in a non-instructional week is removed from the streak calculation entirely, so a learner whose only activity lands in a holiday week has activity_days >= 1 and longest_streak_weeks = 0. That is what a streak over instructional weeks means, but it is why the two columns can disagree and why the checks must expect a zero streak alongside non-zero activity rather than assert a floor of 1.
- Apply the island trick on the re-indexed sequence: instructional_index - ROW_NUMBER() OVER (PARTITION BY learner_id, term_id ORDER BY instructional_index) is constant within a run. Group on that constant, count rows per group, take the max per learner and term.
- Build activity_days from the unfiltered term aggregate, then LEFT JOIN the streak result onto it and COALESCE longest_streak_weeks to 0. An inner join here deletes the holiday-only learners from the output and from the return-rate denominator, which is the same bug as the streak-floor assertion wearing different clothes.
- Set the activity floor on COUNT(DISTINCT activity_date) >= 5 across term T, applied to the return-rate denominator only. The floor exists to remove provisioned-but-unused accounts, so leaking it into the numerator silently redefines the metric as active in both terms. Note that a learner can clear the floor on holiday-week activity alone and enter the denominator with a zero streak.
- Define the return flag as EXISTS any activity row in term T+1 for that learner, pairing terms by an explicit term ordering rather than date arithmetic, since term lengths differ. Learners whose org has no term T+1 loaded must be excluded and counted, not treated as non-returners.
Worked solution 40 min
- Build the learner-week activity CTE with COUNT(DISTINCT activity_date) and confirm one row per (learner_id, term_id, term_week_no).
- Build the instructional re-index from dim_term_week and join activity onto it, dropping non-instructional weeks entirely.
- Apply the index-minus-row-number island grouping and take MAX(run_length) per learner and term.
- Compute activity_days per learner and term over the whole term, LEFT JOIN the streak onto it with COALESCE to 0, and count the learners who come out with activity_days > 0 and longest_streak_weeks = 0.
- Apply the five-day floor to the denominator and attach the term T+1 existence flag.
- Build a two-row fixture by hand and verify the holiday behaviour before trusting the full run.
Follow-up
- This streak rewards a learner doing one minute a week over one doing four hours in a single week. Which of those is the product claim, and what would you pair the streak with to catch the difference?
- Roster sync deactivates accounts in bulk at term end. How does that interact with the five-day floor and with the return denominator?
- Show how return rate moves as the floor goes from one day to five, and argue which floor is the honest one to publish.
How do you prioritize tasks when managing multiple projects?
How do you prioritize tasks when managing multiple projects?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
Describe your method for evaluating the effectiveness of a new educati…
Describe your method for evaluating the effectiveness of a new educational program using data.
Approach
- Fix the population and the time window before naming any metric.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
If given a dataset, how would you identify patterns or anomalies?
If given a dataset, how would you identify patterns or anomalies?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How would you approach a project to predict student enrollment trends?
How would you approach a project to predict student enrollment trends?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Walk us through your thought process in designing an experiment to tes…
Walk us through your thought process in designing an experiment to test a hypothesis.
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
Explain the difference between supervised and unsupervised learning.
Explain the difference between supervised and unsupervised learning.
Approach
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer would contain.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Decide what an administrator usage report counts
You must publish a five-line usage report that institutional administrators read before renewal. Once it is published, instructors will assign work to move whatever it counts. Available: fct_lesson_activity (active_seconds, is_assigned, completion_status), fct_assessment_response (max_points, scoring_mode, is_correct, response_seconds, hint_count), fct_enrollment (units_completed, units_total) and dim_learner (first_activity_at, age_gated). Choose the five lines and their source columns, name one metric you refuse to publish and the behaviour it would cause, and state the pairing that keeps each line honest.
Approach
- Start from behaviour rather than from data availability. Publishing a number to the buyer converts it into an assignment, so the test for each line is whether its gamed version is still something you want happening.
- Choose lines that survive that test: seat activation with the never-activated count shown rather than hidden; learners with at least one scored submission in the last four weeks; pace adherence at the current term-week; verified mastery events with the reporting lag printed; and objectives with no coverage yet, which is a gap list, is directly actionable, and cannot be inflated.
- Name the source columns per line so the report is reproducible: first_activity_at against seats_provisioned, fct_assessment_response.submitted_at with scoring_mode, units_completed over units_total against the term dimension, the mastery ledger with its delayed check, and the objective coverage join.
- Refuse active minutes per learner. Publishing it causes seat-time assignments, which is the clearest case of teaching to the number in this domain and the weakest claim the product can defend as learning. Keep it internally as a guardrail where it is useful and harmless.
- Pair every countable line with an integrity partner: pace adherence with assessment integrity (response_seconds above the item floor, hint_count below cap), submissions with mean max_points so splitting one task into five cannot move the line, activation with the raw never-activated count.
- State the population on the report itself, including the age_gated share and any excluded enrollment sources, because those learners are missing from the behavioural tables and the report will otherwise look complete while describing a partial population.
Worked solution 30 min
- For each candidate line, write the one sentence describing how an instructor would move it with no learning, and drop the lines whose gamed version is undesirable.
- Write the five surviving lines with their source columns and windows.
- Attach an integrity or substance partner to each countable line, and write the single sentence justifying the refusal of active minutes per learner.
- Add the population footer: age_gated share, excluded enrollment sources, and the reporting lag on verified mastery.
Follow-up
- An administrator says your report shows fewer active learners than their own dashboard. What are the three likeliest causes?
- A district asks for a per-teacher leaderboard built from these lines. What do you do?
- If you could publish only four lines, which one goes, and what breaks?
Scored submissions sag in the dashboard's final five days
A daily chart counts fct_assessment_response rows by DATE(scored_at). The last five days each sit below the one before, and every Saturday and Sunday is roughly half of a weekday. Columns available: response_id, learner_id, submitted_at, scored_at, scoring_mode, max_points, attempt_no. Say whether scored submissions are actually declining, produce the corrected series, and state the reporting lag the correction forces on this dashboard.
Approach
- Rebuild the same count on DATE(submitted_at) over sixty days and overlay the two. scored_at lags submitted_at for rubric-scored work, so a scored_at series is right-censored at the recent end: the last days are unfinished, not falling.
- Quantify the lag rather than asserting it. Compute percentiles of scored_at minus submitted_at split by scoring_mode; auto sits near zero and rubric_human carries the tail that produces the sag.
- Take the rubric_human p95 lag as the close threshold and truncate or grey out the scored_at series at max(date) minus that threshold, so the chart cannot be read as a decline again.
- Handle the weekly pattern separately from the lag: compare each day to the same weekday a week earlier, or report a trailing seven-day sum. In an academic product weekend volume is structurally a fraction of a weekday's and is not signal.
- Restate the dashboard definition with the lag written on it, and name which of the two series answers which question.
Follow-up
- What reporting lag would you publish, and what do you do when a rubric grading backlog makes the lag itself drift?
- Which series belongs on an operational grading-capacity dashboard, and why is the other one the right basis for the learning metric?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Data people depend on systems owned by other teams, and much of the job is negotiating for instrumentation, access, or a fix to a broken pipeline. Prepare an example of getting something changed upstream that you did not control. Describe what you asked for, what you traded, and how you worked while you waited.
Describe a situation where you had to work with a difficult team membe…
Describe a situation where you had to work with a difficult team member. How did you handle it?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Pick a story where you drove the decision, not one where you observed it.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that project again?
Account for an analysis you later discovered was wrong
Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.
Approach
- Pick an error with a nameable mechanism, such as a grain mistake, a survivorship filter, or a join that silently dropped age_gated learners. 'I misread a chart' has nothing testable in it.
- State the blast radius in decisions rather than in dashboards: what was decided, what was spent or promised, and to whom.
- Describe detection honestly, including whether someone else caught it and how long it ran undetected. A slow external catch is a usable answer; a vague one is not.
- Give the disclosure order and the reason for it: the decision owner first, with the corrected number and its direction in the opening sentence, then the wider audience.
- Close on the systemic fix and say whether it has fired since. A control that has never triggered on anything is a resolution, not a fix, and saying so is worth more than claiming otherwise.
Follow-up
- What did the corrected number change about the decision? If nothing changed, why did you escalate at all?
- Has the check you added caught anything since, and on what?
- What class of error would still get past you today?
Explain a risk score to a non-technical sales executive
Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.
Approach
- Answer the question that was asked before qualifying it. 'Treat this as at-risk and act this week' is the answer; the probability is the support for it.
- Translate the score into a frequency on a reference class the VP already trusts, such as how many of the last fifty accounts scored in the 0.55-0.70 band did not renew, instead of explaining calibration in the abstract.
- Name the driver that is still movable before period_end: 31% seat activation at term-week 6 can be changed this term, whereas last term's usage cannot, so that is where the conversation goes.
- State what would move the score and by when, so the number arrives attached to an action rather than as a verdict.
- For the written line, use a band and a review date rather than a bare decimal, because a single number in a document will be re-quoted without its interval.
Follow-up
- The VP asks for every account scored above 0.50. What do you tell them about where that threshold came from?
- How would you check whether 0.62 is calibrated at all, and on what sample?
- 01
Describe a situation where you had to work with a difficult team member. How did you handle it?
- 02
Describe a case where you reported a result and later found it was wrong. Choose one where the error was yours and the number had already been used for a decision. Cover: the claim, the actual defect at the data or design level, how you found it, how long it had been live, who you told and in what order, and what you changed so the same class of defect cannot recur. The interviewer will push on the mechanism, not the apology. Probed: whether you understand your own failure modes well enough to have engineered around them.
- 03
Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.
Is this an official University of Michigan interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at University of Michigan. Rounds and questions reflect what candidates have reported, not a process University of Michigan has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process for a Data Scientist position at the University of Michigan?
The interview process is considered rigorous, requiring a solid understanding of technical concepts, problem-solving abilities, and effective communication skills. Candidates should prepare thoroughly to demonstrate their capabilities across these areas.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates typically showcase a strong technical foundation, the ability to communicate insights effectively, and a collaborative mindset. They also align well with the university's values and demonstrate a passion for data-driven decision-making.
PracHub interview research ↗What is the typical timeline from the initial screen to an offer?
The timeline may vary, but candidates can generally expect several weeks from the initial screening call to the final decision. It is advisable to remain patient and proactive during this time.
PracHub interview research ↗What is the working culture like at the University of Michigan?
The culture emphasizes collaboration, innovation, and a commitment to excellence in research and education. Teamwork is highly valued, and there is a strong focus on leveraging data to enhance educational and operational outcomes.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22