University of Maryland, Baltimore · Data Scientist
Updated · 2026-09-24

University of Maryland, Baltimore Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist within the Anesthesiology Faculty at the University of Maryland, Baltimore (UMB), you occupy a critical intersection between advanced clinical research and quantitative analysis. This role is not merely about building models; it is about driving evidence-based breakthroughs in medical care. You will serve as a primary analytical partner for faculty and researchers, translating complex medical datasets into actionable insights that can influence clinical outcomes and institutional research strategies.

Coding rounds for this role are usually data-manipulation shaped rather than data-structure shaped: group-bys, joins, time windows, ranking within a partition. Confirm the format before spending a week on graph traversal.

PracHub has no confirmed round sequence for University of Maryland, Baltimore. Treat the sections below as preparation areas and confirm the format with your recruiter.

Cluster experiments at class or school levelSeparate mastery signals from raw usage exposureReconstruct active time from raw heartbeat streams

25 min read

Practice 11 Data Scientist prompts
11Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist within the Anesthesiology Faculty at the University of Maryland, Baltimore (UMB), you occupy a critical intersection between advanced clinical research and quantitative analysis. This role is not merely about building models; it is about driving evidence-based breakthroughs in medical care. You will serve as a primary analytical partner for faculty and researchers, translating complex medical datasets into actionable insights that can influence clinical outcomes and institutional research strategies.

The work is characterized by high complexity and the need for rigorous statistical integrity. You will navigate diverse healthcare datasets, likely involving perioperative records, patient outcomes, and clinical trial results. Your impact is direct: your findings contribute to the academic mission of UMB, helping to refine anesthetic protocols and improve patient safety. Success in this role requires a candidate who is comfortable operating in an academic research environment, where precision, peer-review-ready documentation, and collaborative problem-solving are paramount.

01

Preparation focus

editorial

No round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.

What to demonstrate

  • Breadth across SQL, experimentation and product reasoning
  • Ability to state assumptions before choosing a method

How to prepare

  • Drill the practice exercises below and time yourself
  • Prepare three quantified stories about decisions you drove
PracHub interview preparation framework ↗

PracHub editorial advice for the preparation topics above.

01

Consent and age rules silently truncate the data, not just the joins.

Where age_gated is TRUE, behavioural logging and cross-system joins are restricted, so those learners are missing from the very tables used to compute engagement. An analysis that simply inner-joins will produce a population skewed to older learners and self-serve accounts while appearing complete. Check the age_gated share of every population you report on, and state it.

02

Adaptive item selection holds observed accuracy flat by construction.

A selector targeting a fixed success probability, say 70%, will keep measured first-attempt accuracy near 70% whether learners are improving or not, because as theta rises the engine simply serves harder items. Reporting 'accuracy improved 3 points' on adaptive content therefore usually means the selector got more conservative, not that anyone learned more. Accuracy is only interpretable on a fixed form with unchanged item versions, which is why the first-attempt accuracy metric above carries both restrictions.

03

Extrapolating a first-week lift inflated by novelty effects

Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.

04

Writing SQL without stating NULL and tie-breaking behaviour

Before calling a query finished, say what it does with NULLs, ties and empty groups. NOT IN against a subquery containing a single NULL returns no rows at all, and RANK, DENSE_RANK and ROW_NUMBER differ precisely on ties, so name which one the question requires.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

8 technical prompts3 include a worked solution

Explain the difference between fixed-effects and random-effects models…

medium
statistics and probability

Explain the difference between fixed-effects and random-effects models in the context of longitudinal patient data.

Approach
  1. Write down the assumption the method needs before you use the method.
  2. Sanity-check the answer against a simple bound or a simulated case.
  3. Say what the estimate is of, and over what population it generalises.
Follow-up
  • How would you explain this result to someone who does not know statistics?
  • What sample size would you need to detect an effect half this size?

Attach the calibration in force without merge_asof

medium
as-of joinsearchsortedvectorisationslowly changing dimensions

You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.

Approach
  1. Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
  2. Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
  3. Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
  4. Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
  5. Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
  6. Take the parameters positionally with np.take and write NaN where any invalidation fired.
Follow-up
  • A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
  • The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
  • How would you test this without a golden output to compare against?

Sessionise a raw heartbeat stream into activity rows

hardWorked solution
sessionisationevent streamsgroupbytime series

You get a DataFrame heartbeats with columns learner_id, content_item_id, content_version_no, and ts (UTC timestamps, unsorted). Produce one row per learner per session per content item with started_at, ended_at, wall_seconds, and active_seconds. A session closes after 30 minutes with no heartbeat from that learner. active_seconds sums consecutive gaps inside a session, capping each gap at 120 seconds. Attribute each gap to the content item of the earlier heartbeat. State your rule for a gap of exactly 1800 seconds and for a session with a single heartbeat.

Approach
  1. Sort by learner_id then ts, and compute the gap with a per-learner diff so the first heartbeat of each learner has no gap. A global diff lets one learner's last heartbeat set the next learner's first gap.
  2. Flag a session break where gap > 1800 seconds, then cumsum that boolean within the learner to get a session index. Say whether exactly 1800 opens a new session; either choice is defensible but it must be stated because it moves the session count.
  3. Each gap covers the interval before the current heartbeat, so attribute min(gap, 120) to the content_item_id of the previous row via shift, not the current row.
  4. Group by (learner_id, session_index, content_item_id): sum the attributed capped gaps for active_seconds, take min and max ts for started_at and ended_at, and derive wall_seconds from those two.
  5. Handle the degenerate group explicitly. A single heartbeat gives active_seconds 0 and wall_seconds 0. Dropping those rows quietly removes real opens and biases any completion or abandonment rate computed later.
  6. Validate that within every session the summed active_seconds never exceeds the session's wall_seconds, which catches attribution and cap errors in one assertion.
Worked solution 35 min
  1. Sort by (learner_id, ts) and compute gap = ts.diff().dt.total_seconds(), then null the gap on the first row of each learner using a learner-change mask.
  2. new_session = gap.isna() | (gap > 1800); session_index = new_session.groupby(learner_id).cumsum().
  3. contribution = gap.clip(upper=120).fillna(0), attributed to the previous row's content_item_id using shift within the learner and session.
  4. Aggregate by (learner_id, session_index, content_item_id) with sum for active_seconds and min and max of ts for the timestamps.
  5. Assert per session that sum(active_seconds) <= (max(ts) - min(ts)) in seconds, and print the count of single-heartbeat groups.
EXPECTED RESULTA frame keyed by (learner_id, session_index, content_item_id) where active_seconds equals the sum of min(gap, 120) over consecutive heartbeats with the first heartbeat contributing zero, wall_seconds equals ended_at minus started_at, and sum(active_seconds) <= wall_seconds holds for every session.
Follow-up
  • The client emits a heartbeat every 60 seconds on desktop and every 15 seconds on a managed device fleet. What does the 120-second cap do to the comparison between those two populations?
  • Your reconstructed active_seconds disagrees with the stored column by 4 percent overall. How do you localise the disagreement rather than argue about the total?
  • A learner has two tabs open on different items at once. What does your output claim, and what should it claim?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Saying no well is a senior skill and it is rarely rehearsed. Think of a time you told someone their analysis was not worth doing, or that the experiment could not answer their question at the sample size available. Explain what you offered instead. Refusal without an alternative reads as obstruction rather than judgement.

How do you handle missing or incomplete data in clinical research sets…

medium
behavioural and stakeholder questions

How do you handle missing or incomplete data in clinical research sets?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. State the situation in two sentences and spend the rest on your reasoning.
  3. Quantify the outcome, including what you would not claim credit for.
Follow-up
  • What did you decide not to do, and why?
  • What would you do differently if you ran that project again?

Describe your experience using SAS for complex data manipulation and s…

medium
behavioural and stakeholder questions

Describe your experience using SAS for complex data manipulation and statistical modeling.

Approach
  1. State the situation in two sentences and spend the rest on your reasoning.
  2. Pick a story where you drove the decision, not one where you observed it.
  3. Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
  • What would you do differently if you ran that project again?
  • How did you know the outcome was caused by your change?

Explain a risk score to a non-technical sales executive

easy
communicating uncertaintyrenewal riskstakeholder communication

Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.

Approach
  1. Answer the question that was asked before qualifying it. 'Treat this as at-risk and act this week' is the answer; the probability is the support for it.
  2. Translate the score into a frequency on a reference class the VP already trusts, such as how many of the last fifty accounts scored in the 0.55-0.70 band did not renew, instead of explaining calibration in the abstract.
  3. Name the driver that is still movable before period_end: 31% seat activation at term-week 6 can be changed this term, whereas last term's usage cannot, so that is where the conversation goes.
  4. State what would move the score and by when, so the number arrives attached to an action rather than as a verdict.
  5. For the written line, use a band and a review date rather than a bare decimal, because a single number in a document will be re-quoted without its interval.
Follow-up
  • The VP asks for every account scored above 0.50. What do you tell them about where that threshold came from?
  • How would you check whether 0.62 is calibrated at all, and on what sample?
  • 01

    How do you handle missing or incomplete data in clinical research sets?

  • 02

    Describe your experience using SAS for complex data manipulation and statistical modeling.

  • 03

    Your renewal-risk score joins fct_subscription_period to seat activation from dim_learner.first_activity_at and to pace adherence from fct_enrollment. For one institution with 900 seats_provisioned and 31% activation at term-week 6, it outputs a 0.62 probability of non-renewal. A sales VP asks whether you are losing the account or not. You have two minutes and no slides. Deliverable: the spoken answer, plus the one line you would put in the weekly account review so the number is not later quoted as a certainty. Probed: whether you can carry uncertainty and still give a decision.

PracHub interview preparation framework ↗
Is this an official University of Maryland, Baltimore interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at University of Maryland, Baltimore. Rounds and questions reflect what candidates have reported, not a process University of Maryland, Baltimore has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult is the technical assessment?

The assessment is designed to be practical. It is less about "trick" algorithm questions and more about your ability to perform real-world data manipulation and statistical tasks relevant to the role.

PracHub interview research ↗
What is the team culture like?

It is an academic, research-driven environment. You will be working with experts, so expect a culture that values intellectual curiosity, precision, and collaborative feedback.

PracHub interview research ↗
Will I have to present my work?

Yes. You may be asked to explain your approach to a previous project or present the results of your technical assessment to the research team.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.