As a Data Scientist within the Anesthesiology Faculty at the University of Maryland, Baltimore (UMB), you occupy a critical intersection between advanced clinical research and quantitative analysis. This role is not merely about building models; it is about driving evidence-based breakthroughs in medical care. You will serve as a primary analytical partner for faculty and researchers, translating complex medical datasets into actionable insights that can influence clinical outcomes and institutional research strategies.
The work is characterized by high complexity and the need for rigorous statistical integrity. You will navigate diverse healthcare datasets, likely involving perioperative records, patient outcomes, and clinical trial results. Your impact is direct: your findings contribute to the academic mission of UMB, helping to refine anesthetic protocols and improve patient safety. Success in this role requires a candidate who is comfortable operating in an academic research environment, where precision, peer-review-ready documentation, and collaborative problem-solving are paramount.
Preparation focus
editorialNo round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.
What to demonstrate
- Breadth across SQL, experimentation and product reasoning
- Ability to state assumptions before choosing a method
How to prepare
- Drill the practice exercises below and time yourself
- Prepare three quantified stories about decisions you drove
PracHub editorial advice for the preparation topics above.
Consent and age rules silently truncate the data, not just the joins.
Where age_gated is TRUE, behavioural logging and cross-system joins are restricted, so those learners are missing from the very tables used to compute engagement. An analysis that simply inner-joins will produce a population skewed to older learners and self-serve accounts while appearing complete. Check the age_gated share of every population you report on, and state it.
Adaptive item selection holds observed accuracy flat by construction.
A selector targeting a fixed success probability, say 70%, will keep measured first-attempt accuracy near 70% whether learners are improving or not, because as theta rises the engine simply serves harder items. Reporting 'accuracy improved 3 points' on adaptive content therefore usually means the selector got more conservative, not that anyone learned more. Accuracy is only interpretable on a fixed form with unchanged item versions, which is why the first-attempt accuracy metric above carries both restrictions.
Extrapolating a first-week lift inflated by novelty effects
Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.
Writing SQL without stating NULL and tie-breaking behaviour
Before calling a query finished, say what it does with NULLs, ties and empty groups. NOT IN against a subquery containing a single NULL returns no rows at all, and RANK, DENSE_RANK and ROW_NUMBER differ precisely on ties, so name which one the question requires.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the difference between fixed-effects and random-effects models…
Explain the difference between fixed-effects and random-effects models in the context of longitudinal patient data.
Approach
- Write down the assumption the method needs before you use the method.
- Sanity-check the answer against a simple bound or a simulated case.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
Attach the calibration in force without merge_asof
You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.
Approach
- Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
- Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
- Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
- Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
- Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
- Take the parameters positionally with np.take and write NaN where any invalidation fired.
Follow-up
- A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
- The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
- How would you test this without a golden output to compare against?
Sessionise a raw heartbeat stream into activity rows
You get a DataFrame heartbeats with columns learner_id, content_item_id, content_version_no, and ts (UTC timestamps, unsorted). Produce one row per learner per session per content item with started_at, ended_at, wall_seconds, and active_seconds. A session closes after 30 minutes with no heartbeat from that learner. active_seconds sums consecutive gaps inside a session, capping each gap at 120 seconds. Attribute each gap to the content item of the earlier heartbeat. State your rule for a gap of exactly 1800 seconds and for a session with a single heartbeat.
Approach
- Sort by learner_id then ts, and compute the gap with a per-learner diff so the first heartbeat of each learner has no gap. A global diff lets one learner's last heartbeat set the next learner's first gap.
- Flag a session break where gap > 1800 seconds, then cumsum that boolean within the learner to get a session index. Say whether exactly 1800 opens a new session; either choice is defensible but it must be stated because it moves the session count.
- Each gap covers the interval before the current heartbeat, so attribute min(gap, 120) to the content_item_id of the previous row via shift, not the current row.
- Group by (learner_id, session_index, content_item_id): sum the attributed capped gaps for active_seconds, take min and max ts for started_at and ended_at, and derive wall_seconds from those two.
- Handle the degenerate group explicitly. A single heartbeat gives active_seconds 0 and wall_seconds 0. Dropping those rows quietly removes real opens and biases any completion or abandonment rate computed later.
- Validate that within every session the summed active_seconds never exceeds the session's wall_seconds, which catches attribution and cap errors in one assertion.
Worked solution 35 min
- Sort by (learner_id, ts) and compute gap = ts.diff().dt.total_seconds(), then null the gap on the first row of each learner using a learner-change mask.
- new_session = gap.isna() | (gap > 1800); session_index = new_session.groupby(learner_id).cumsum().
- contribution = gap.clip(upper=120).fillna(0), attributed to the previous row's content_item_id using shift within the learner and session.
- Aggregate by (learner_id, session_index, content_item_id) with sum for active_seconds and min and max of ts for the timestamps.
- Assert per session that sum(active_seconds) <= (max(ts) - min(ts)) in seconds, and print the count of single-heartbeat groups.
Follow-up
- The client emits a heartbeat every 60 seconds on desktop and every 15 seconds on a managed device fleet. What does the 120-second cap do to the comparison between those two populations?
- Your reconstructed active_seconds disagrees with the stored column by 4 percent overall. How do you localise the disagreement rather than argue about the total?
- A learner has two tabs open on different items at once. What does your output claim, and what should it claim?
Can you walk us through a time you had to select a statistical test fo…
Can you walk us through a time you had to select a statistical test for a non-normal distribution?
Approach
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Find enrollments with no scored work using anti-joins
You have fct_enrollment(enrollment_id, learner_id, course_id, section_id, org_id, status, enrolled_at, due_date) and fct_assessment_response(response_id, learner_id, enrollment_id, submitted_at, scored_at), where fct_assessment_response.enrollment_id is NULL for work done outside any enrollment. Return every active enrollment with a due_date in the next 14 days that has no scored response attached, for an instructor at-risk list. Write it as an anti-join. A colleague's NOT IN version returns zero rows against production but the correct count against their small test extract. Explain exactly why, and fix it.
Approach
- Name the mechanism precisely: NOT IN expands to a conjunction of inequality comparisons, and x <> NULL is unknown rather than false, so once a single NULL enters the subquery list the predicate can never evaluate to TRUE and the result is empty. The test extract contained no work outside an enrollment, so it contained no NULLs, so it passed.
- Fix with NOT EXISTS and a correlated equality on enrollment_id. It tests row existence rather than list membership and is unaffected by NULLs elsewhere in the column. LEFT JOIN with an IS NULL filter is equivalent and usually plans the same; NOT IN with an added IS NOT NULL guard also works but leaves the landmine armed for the next reader.
- Decide what no scored work means before writing it. No response at all and responses that exist but are unscored are different at-risk lists, and rubric lag puts already-submitted work in the second. Return both counts rather than collapsing them.
- Filter status = 'active' and due_date BETWEEN CURRENT_DATE AND CURRENT_DATE + 14, comparing DATE to DATE. Casting a DATE against a timestamp midnight quietly drops the final day.
- Guard the complement in the other direction: joining enrollments to responses and counting enrollments fans out one enrollment with 40 responses into 40 rows, so the has-scored-work list needs COUNT(DISTINCT enrollment_id).
Worked solution 20 min
- Run SELECT COUNT(*) FROM fct_assessment_response WHERE enrollment_id IS NULL. That one number is the whole explanation.
- Write both versions side by side over the same window and show the row counts differ.
- Write the NOT EXISTS version with the response_count and scored_response_count columns attached.
- Verify the partition of the due-date window: at-risk count plus distinct enrollments with scored work equals total active enrollments in the window.
Follow-up
- Produce the complement list, the enrollments that do have scored work, without double counting, and state the grain of your join.
- An enrollment whose learner submitted work with enrollment_id NULL has done the work but is not attached to it. How would you attribute those rows, and what does attributing by learner and course risk?
- fct_assessment_response is partitioned by date. How would you bound the scan without changing the answer?
How do you explain complex statistical findings to a PI or researcher …
How do you explain complex statistical findings to a PI or researcher who does not have a background in data science?
Approach
- Work from the decision backwards to the evidence you would need.
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Apply CUPED and stratification to an under-powered clustered test
A section-randomised test is under-powered and the team cannot get more sections. Eight term-weeks of pre-period fct_assessment_response and fct_lesson_activity exist before assignment. The learner-level correlation between pre-period and in-test weekly mastery events is 0.55; at section-mean level it is 0.72. Propose a variance-reduction plan. Compute the MDE multiplier CUPED buys at each level, say which level you apply it at and why, and handle the learners with no pre-period because roster_sync created their accounts after assignment.
Approach
- State the estimator: Y_adj = Y - theta x (X - mean(X)), where X is the pre-period covariate and the binding precondition is that X is measured strictly before assignment. Fit theta = Cov(Y, X) / Var(X) on the in-test data pooled across both arms. That is standard CUPED and it is the only option here: Y is the in-test outcome, so theta cannot be computed from pre-period data alone. Pooling is what keeps it honest, because a pre-assignment X has the same distribution in both arms and the adjustment therefore cannot absorb the treatment effect. Var(Y_adj) = (1 - r^2) x Var(Y), so the MDE scales by sqrt(1 - r^2).
- Apply it at the unit of analysis. Under section randomisation that is the section mean, so the 0.72 section-level correlation is the one that matters. It exceeds the learner-level 0.55 because averaging removes idiosyncratic learner noise while preserving the section's persistent level.
- Stratify the randomisation itself rather than only adjusting afterwards: block sections on pre-period mastery-event quartile crossed with grade_band, then randomise within block. With 300 sections per arm, blocking also protects against a bad draw, which post-hoc CUPED cannot undo.
- Handle the no-pre-period learners without imputing a value that would drag theta. Place them in their own stratum with X set to that stratum's mean, which leaves them with zero adjustment and no bias, and report their share alongside the result.
- Hold the two rules that actually protect the estimate: the covariate must be measured strictly before assignment, and theta must be fitted once on the pooled arms rather than separately within each arm. A post-assignment covariate can itself be moved by treatment, so adjusting on it subtracts part of the effect you are measuring. Arm-specific thetas make the adjustment differ by arm, which reintroduces the same bias by another route. Fitting theta on the data it adjusts carries a finite-sample bias of order 1 / n_clusters; at 300 sections per arm that is negligible, and if you want it gone entirely, estimate theta from an earlier pre-period regression on data outside the test window and accept a slightly lower r.
Worked solution 30 min
- Learner level: sqrt(1 - 0.55^2) = sqrt(0.6975) = 0.835, a 16.5% smaller MDE.
- Section level: sqrt(1 - 0.72^2) = sqrt(0.4816) = 0.694, a 30.6% smaller MDE.
- Convert to sample equivalence: 1 / 0.4816 = 2.08, so section-level CUPED is worth roughly doubling the number of sections.
- Build the blocks as four pre-period quartiles crossed with grade_band, randomise within block, and persist the block id for the analysis.
- Fit theta once on the pooled section-level in-test data, form Y_adj, then estimate the effect at section level with block fixed effects and compare the adjusted and unadjusted estimates and intervals.
Follow-up
- Your pre-period window spans a holiday for half the sections. What does that do to the correlation and to theta?
- Can you stack CUPED on top of stratification, and does the second one still buy anything once the first has run?
- The section-level correlation is 0.72 overall but 0.45 for sections whose instructor changed between periods. How do you handle that subgroup?
Net revenue retention slid nine points in one quarter
Net revenue retention on institutional licenses fell from 104 percent to 95 percent quarter over quarter. Source: fct_subscription_period (subscription_period_id, account_id, org_id, plan_tier, billing_interval, seats_purchased, seats_provisioned, mrr_usd, period_start, period_end, is_renewal, status, canceled_at, churn_reason). Decompose the nine points into expansion, contraction and non-renewal, and establish whether renewal behaviour changed at all. Deliverable: the decomposition, the cohort definition you used, and a call on whether this is a trend or account news.
Approach
- Pin the cohort in writing before computing anything: orgs holding an active period twelve months before each renewal date, new logos excluded from both numerator and denominator, and non-renewals entering the numerator as zero rather than being dropped. Dropping non-renewals is the most common way this metric is overstated, and correcting it can move the headline by more than the effect under investigation.
- Rebuild the ratio from components on that cohort: starting annualised revenue, expansion, contraction, non-renewal. Check the components sum back to the reported ratio in both quarters; a residual means the two quarters were not computed on the same cohort rule and the comparison is void before any interpretation.
- Test cohort comparability on billing_interval. Multi-year contracts do not present a renewal in every window, so a quarter whose renewal cohort holds a different multi_year share is not comparable to its predecessor. Recompute both quarters holding the interval mix fixed and report that alongside the raw number.
- Rank per-org contributions to the nine points and state how much the top three carry. Institutional revenue is concentrated, so a decomposition that leaves the concentration unstated invites a trend reading of what is account news.
- Audit the inputs: confirm mrr_usd normalisation is consistent (annual divided by twelve) and look for overlapping period_start and period_end within a single account_id, which double counts revenue after a mid-term upgrade.
- Close the loop to something actionable by joining the lost orgs back to seat activation and pace adherence in the preceding term, and say whether the signal was visible early enough for anyone to have acted on it.
Follow-up
- Build the renewal-risk view from this: which signals are actionable while there is still time to act, and at what days-to-renewal would you surface an org?
- An org contracts from 900 seats to 600 but moves to a higher tier and raises MRR. Expansion or contraction, and does your decomposition handle it without double counting?
- How would you present NRR so that a quarter with an unusual renewal cohort is not read as a trend by an executive audience?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Saying no well is a senior skill and it is rarely rehearsed. Think of a time you told someone their analysis was not worth doing, or that the experiment could not answer their question at the sample size available. Explain what you offered instead. Refusal without an alternative reads as obstruction rather than judgement.
How do you handle missing or incomplete data in clinical research sets…
How do you handle missing or incomplete data in clinical research sets?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that project again?
Describe your experience using SAS for complex data manipulation and s…
Describe your experience using SAS for complex data manipulation and statistical modeling.
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Explain a risk score to a non-technical sales executive
Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.
Approach
- Answer the question that was asked before qualifying it. 'Treat this as at-risk and act this week' is the answer; the probability is the support for it.
- Translate the score into a frequency on a reference class the VP already trusts, such as how many of the last fifty accounts scored in the 0.55-0.70 band did not renew, instead of explaining calibration in the abstract.
- Name the driver that is still movable before period_end: 31% seat activation at term-week 6 can be changed this term, whereas last term's usage cannot, so that is where the conversation goes.
- State what would move the score and by when, so the number arrives attached to an action rather than as a verdict.
- For the written line, use a band and a review date rather than a bare decimal, because a single number in a document will be re-quoted without its interval.
Follow-up
- The VP asks for every account scored above 0.50. What do you tell them about where that threshold came from?
- How would you check whether 0.62 is calibrated at all, and on what sample?
- 01
How do you handle missing or incomplete data in clinical research sets?
- 02
Describe your experience using SAS for complex data manipulation and statistical modeling.
- 03
Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.
Is this an official University of Maryland, Baltimore interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at University of Maryland, Baltimore. Rounds and questions reflect what candidates have reported, not a process University of Maryland, Baltimore has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the technical assessment?
The assessment is designed to be practical. It is less about "trick" algorithm questions and more about your ability to perform real-world data manipulation and statistical tasks relevant to the role.
PracHub interview research ↗What is the team culture like?
It is an academic, research-driven environment. You will be working with experts, so expect a culture that values intellectual curiosity, precision, and collaborative feedback.
PracHub interview research ↗Will I have to present my work?
Yes. You may be asked to explain your approach to a previous project or present the results of your technical assessment to the research team.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22