As a Data Scientist at The University of Pennsylvania, you occupy a unique position at the intersection of rigorous academic inquiry and high-stakes institutional operations. This role is not merely about running models; it is about providing the quantitative backbone for major initiatives that influence research outcomes, operational efficiency, and the broader university experience. You will be responsible for translating complex, messy datasets into clear, actionable insights that help leadership make evidence-based decisions.
The work is intellectually demanding and requires a high degree of versatility. You will likely collaborate with diverse stakeholders—ranging from department directors to technical engineering teams—to design product metrics, evaluate the success of institutional programs, and maintain the integrity of data pipelines. Whether you are performing a deep-dive analysis on a metric drop or architecting an A/B test to improve user engagement with university digital platforms, your impact is measured by your ability to bridge the gap between raw data and strategic direction.
Because you are working in an academic-institutional setting, be prepared to explain your technical findings to non-technical stakeholders who prioritize reliability and long-term institutional impact over pure speed or novelty.
Recruiter Screen
reportedMost candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.
What to demonstrate
- Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
- Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
- The substance of the questions you ask back, which an experienced screener reads as a level signal
How to prepare
- Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
- Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
- Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
Technical Screening
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Behavioral Assessment
reportedRounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.
What to demonstrate
- Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
- What you changed afterwards, stated as a check you now run rather than a lesson you now believe
- Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you
How to prepare
- Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
- Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
- Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
Panel Interview
reportedWhere a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.
What to demonstrate
- Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
- Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
- Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
- Whether you ask what decision is actually on the table before explaining anything
How to prepare
- Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
- For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
- Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub editorial advice for the preparation topics above.
Learner-level standard errors on class-level interventions.
Anything an instructor controls, and anything deployed by school, is assigned at the section or school level, and outcomes within a section are correlated through the shared instructor, schedule, and device fleet. Treating the learner as the unit of independence understates variance by the design effect 1 + (m - 1) * rho for roughly equal cluster sizes. At a typical section size of 25 and rho of 0.15 that is a factor near 4.6 on variance, which routinely converts a null into a 'significant' result.
The academic calendar creates structural breaks that look like product effects.
Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.
Reading experiment results before checking the arm split
Compare observed arm counts against the intended allocation ratio, not an assumed even split, and set the alarm far below the conventional 0.05: at 0.05 roughly one healthy experiment in twenty trips it, which is why sample-ratio checks usually run at p < 0.001 or stricter. The test's power scales with sample size, so it misses a real diversion on a small experiment and fires on an imbalance too small to move the estimate on a very large one. A flag means go find the assignment or logging fault before reading any outcome, not report a mismatch.
Explaining an aggregate move without decomposing the mix shift
Split the change in the aggregate into within-segment movement and movement in segment weights before you explain it. Every segment's rate can fall while the overall rate rises, purely because volume shifted toward segments that already had higher rates.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Implement verified mastery events with a delayed check
From fct_assessment_response (learner_id, objective_id, content_item_id, submitted_at, attempt_no, is_correct, score_points, max_points, scoring_mode, scored_at, mastery_state_after), count verified mastery events for one calendar month. A mastery event is the first time a (learner_id, objective_id) pair reaches mastery_state_after 'mastered'. It is verified when the earliest scored response on that objective submitted at least 7 days later scores score_points/max_points >= 0.8. Return the monthly count and the verified rate, with the denominator restricted to events that have an eligible later check.
Approach
- Filter to rows with mastery_state_after = 'mastered', then take the earliest submitted_at per (learner_id, objective_id). This is not the same as the first row overall for the pair, since earlier attempts sit in 'learning'.
- Restrict those crossings to the target month in UTC. A pair that first crossed in an earlier month never enters this month's count even if it crosses again after a lapse, so the first-crossing filter must run over all history, not just the month.
- Build the candidate check set by joining the full response table back on (learner_id, objective_id) and keeping rows with submitted_at >= crossing + 7 days, score_points not null, scored_at not null, and max_points > 0.
- Take the earliest eligible candidate per pair, not the best one. Taking the maximum score turns a retention check into an ever-passed-later check, which is a much weaker claim and a materially higher number.
- Mark verified where score_points / max_points >= 0.8. The denominator is crossings with at least one eligible candidate; report the count of crossings with none alongside, because that share moves with the reporting lag and a large share means the month is not ready to publish.
Worked solution 30 min
- mastered = df[df.mastery_state_after == 'mastered']; crossings = mastered.groupby(['learner_id','objective_id'], as_index=False).submitted_at.min().
- month_crossings = crossings[crossings.submitted_at.between(month_start, month_end)].
- Merge the full response frame onto month_crossings on (learner_id, objective_id), keep rows where submitted_at >= crossing + 7 days and score_points.notna() and scored_at.notna() and max_points > 0.
- Sort candidates by submitted_at and take the first per pair with drop_duplicates on the pair key.
- verified = (score_points / max_points) >= 0.8; report verified.sum(), len(candidates), and len(month_crossings) - len(candidates).
Follow-up
- A pair crosses to 'mastered', lapses, and crosses again inside the same month. What does your code count, and what should it count?
- Why does this metric need a 30-day reporting lag, and what shape does the series take if you publish it at 7 days?
- The 0.8 ratio is a product choice. How would you check it actually separates learners rather than just passing almost everyone?
Simulate the gain a bottom-quartile selection manufactures
Observed quiz scores are a true score plus independent noise, both normal, with reliability r equal to true-score variance over observed-score variance. Simulate 200,000 learners, standardise the pretest to mean 0 and SD 1, select the bottom quartile on the pretest, and measure the mean pretest-to-posttest change with no intervention at all. Report the manufactured gain for r in {0.5, 0.7, 0.9, 1.0}, and give the closed form your simulation should reproduce. Use numpy only, no statistical packages.
Approach
- Parameterise by r directly: draw the true score with variance r and each noise term with variance 1 - r, so the observed score has unit variance and reliability exactly r by construction. Draw one true score and two independent noise terms per learner.
- Select on the pretest and only the pretest. Selecting on the true score, or on the average of the two measurements, removes the correlation between selection and the pretest's own noise, and the artefact disappears. That sensitivity is the lesson.
- Compute the mean of posttest minus pretest over the selected group. The true score cancels in the difference, so whatever remains is entirely the difference of two noise draws conditioned on the first being low.
- Check against the closed form. For jointly normal standardised scores, E[posttest | pretest] = r * pretest, so the expected gain is (1 - r) * |E[pretest | bottom quartile]|, and E[pretest | bottom quartile] = -phi(z_0.25) / 0.25 = -1.2711.
- Extend the script with a control group selected by the identical rule from an untreated population and show the difference of differences returns to zero. That is the design fix you would actually propose, not a caveat in a footnote.
Follow-up
- A published case study reports large gains specifically for learners who started in the bottom quartile. What do you ask for before believing any of it?
- The posttest is twice as long as the pretest, so its reliability is higher. Does the manufactured gain grow or shrink, and why?
- Give a design that measures a real effect on exactly this selected population without a randomised control.
Attach the calibration in force without merge_asof
You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.
Approach
- Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
- Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
- Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
- Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
- Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
- Take the parameters positionally with np.take and write NaN where any invalidation fired.
Follow-up
- A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
- The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
- How would you test this without a golden output to compare against?
Given two tables, how would you join them to identify users who perfor…
Given two tables, how would you join them to identify users who performed action A but not action B?
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Explain how you would optimize a slow-running SQL query.
Explain how you would optimize a slow-running SQL query.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Daily scored submissions under rubric scoring lag
fct_assessment_response(response_id, learner_id, content_item_id, content_version_no, submitted_at, is_correct, score_points, max_points, scoring_mode, scored_at) stores both timestamps in UTC. scored_at is NULL until a rubric item is graded and lags submitted_at by days when scoring_mode = 'rubric_human'. Produce two dense daily series over the last 90 days: submissions keyed on submission date, and completed scorings keyed on scored date. Then state the reporting lag after which each series stops moving, computed from the data rather than asserted, and say what biases your lag estimate.
Approach
- Key the two series on different columns on purpose: date(submitted_at) for volume, date(scored_at) for graded throughput, with scored_at IS NOT NULL on the second. They answer different questions and are not two versions of one number.
- Define scored as scored_at IS NOT NULL rather than is_correct IS NOT NULL. Check first whether rubric rows carry score_points with is_correct left NULL after grading; if they do, the is_correct test undercounts by exactly the rubric share.
- Estimate lag empirically with percentiles of scored_at - submitted_at grouped by scoring_mode. This is computed only over rows already scored, so it is survivorship-biased downward. Report alongside it the share of submissions in each age bucket still unscored.
- Generate a dense date spine and LEFT JOIN both series onto it so a zero-volume day is a 0, not an absent row that a chart will interpolate across.
- Name the timezone problem instead of guessing: these tables carry a locale on the learner but no timezone, so a local-calendar day series is not computable from this schema. Say that rather than pretending date(submitted_at) is an institution's day.
Worked solution 15 min
- Build the date spine for the 90-day window, then the two aggregates, then join.
- Run a distribution of is_correct for rows with scored_at IS NOT NULL to decide which column defines scored.
- Compute median and p95 of scored_at - submitted_at per scoring_mode, and the unscored share by 1, 3, 7 and 14 day age buckets.
- Write one sentence stating the reporting lag you are imposing and what it costs.
Follow-up
- The last five days of the submitted series will be revised upward as late rows land. How do you present a series whose recent points are known to be incomplete?
- Rubric grading is done by people with weekly schedules. What does that do to a daily scored series, and would an ISO week grain be more honest?
- If you had to pick a single cut-off date after which the data is considered final, how would you derive it and what percentage of eventual scorings would you be discarding?
You notice a sudden metric drop in daily active users; how do you go a…
You notice a sudden metric drop in daily active users; how do you go about diagnosing the root cause?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design a product metric to measure the success of a new …
How would you design a product metric to measure the success of a new student-facing portal?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How do you balance short-term engagement metrics with long-term user s…
How do you balance short-term engagement metrics with long-term user satisfaction?
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you prioritize tasks when you have competing requests from diff…
How do you prioritize tasks when you have competing requests from different departments?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you determine the statistical significance of an experiment res…
How do you determine the statistical significance of an experiment result?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
What are the risks of performing multiple comparisons in a single expe…
What are the risks of performing multiple comparisons in a single experiment?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
Pick randomisation units for two features with different carryover
Two features need tests. Feature A shows each learner the peer explanations other learners in the same class section submitted. Feature B changes the turnaround target for rubric_model scoring from four hours to fifteen minutes, served from one shared scoring queue. Product wants both learner-randomised. Using fct_lesson_activity.is_assigned, fct_assessment_response.scoring_mode and fct_enrollment.section_id, propose a randomisation unit for each, name the interference each one suffers, and state the condition that makes a switchback design valid for one of them and invalid for the other.
Approach
- Feature A leaks by construction. A treated learner's explanation renders to control learners in the same section, so the control arm is partly treated and the estimate is attenuated toward zero by an amount you cannot recover. Randomise at section_id, or at the school when sections share an instructor or a device cart.
- Price the cluster design before agreeing to it. At section size 25 and intraclass correlation 0.15 the design effect is 1 + 24 x 0.15 = 4.6, so budget roughly 4.6 times the learner count a learner-randomised test would have needed.
- Feature B's interference runs through the shared resource, not the classroom. Serving a fifteen-minute target to half the traffic changes queue depth for the other half, so unit-level randomisation is contaminated at the queue regardless of who the units are.
- Apply the switchback condition: the treatment effect must not persist past the switch. Feature B acts on scoring latency, which resets the moment the config flips, so a time-blocked switchback over the whole queue is valid. Randomise the sequence and analyse at the block level, targeting at least 30 blocks per arm because blocks, not learners, are the unit.
- Feature A fails that condition. Seeing a peer explanation changes what a learner knows, and knowledge does not revert when the flag does, so every post-switch control block is permanently contaminated. Switchback is the wrong tool for anything that changes mastery state.
- Write a washout into Feature B's plan: discard the first block after each switch while the queue drains, and confirm the drain time from the observed scored_at minus submitted_at distribution rather than assuming it.
Worked solution 30 min
- For each feature, write the exposure mapping explicitly: whose outcome changes when one unit is treated.
- For Feature A, set the unit to section_id and compute the sample multiplier 1 + (m - 1) x rho = 4.6 at m = 25, rho = 0.15.
- For Feature B, fix a six-hour block, count the blocks available in the test window, and confirm at least 30 blocks land in each arm.
- Specify analysis up front: section-clustered standard errors for A, block-level means with a randomisation-inference p-value for B.
- Add the washout rule for B and a spillover diagnostic for A comparing control-arm outcomes in sections with high versus low treated share.
Follow-up
- For Feature B, how do you choose the switch interval, and what does shortening it cost you?
- Feature A is section-randomised, but two sections in the study share an instructor who teaches both. What do you do about that pair?
- How would you measure the size of Feature A's spillover rather than only designing around it?
Verified mastery rate fell six points with usage flat
Verified mastery rate fell from 0.72 to 0.66 month over month while active learners and submissions per learner held flat. Tables: fct_assessment_response (learner_id, objective_id, content_item_id, content_version_no, attempt_no, submitted_at, is_correct, mastery_state_after, mastery_prob_after) and dim_content_item (content_item_id, version_no, irt_b, irt_a, calibration_n). The metric carries a 30-day reporting lag and requires a delayed retention check at least 7 days after the mastery crossing. Establish whether learning regressed. Deliverable: a ranked list of causes, each eliminated by a named check.
Approach
- Split the rate into its two counts and chart them separately: crossings with an eligible delayed check available, which is the denominator, and crossings that passed it, which is the numerator. A rate that moves while its denominator is also moving is a composition question before it is a learning question.
- Test censoring next, because it is cheap and it invalidates everything downstream. Recompute the prior month's value as it stood at the same maturity as the current month. If the prior month at equal maturity also read 0.66, there is no regression, only an early read against a fully matured comparison.
- Test the decision rule. mastery_prob_after is a posterior from the knowledge-tracing model, so compare its distribution at the crossing response across the two months rather than only counting crossings. A retrained model or a moved threshold changes which learners enter the numerator at all and shows up as a shifted posterior, not as a usage change.
- Test the instrument. Compute the response-weighted mean irt_b of items used in delayed checks in each month, restricted to items with calibration_n at or above 200, and compare irt_a alongside it. A harder or differently discriminating retention form lowers the pass rate with no change in what learners know.
- Segment only once the artefacts are excluded, and segment cluster-aware: cut by grade_band, locale and is_assigned, then check how concentrated the move is across section_id. Outcomes correlate within sections through a shared instructor, schedule and device fleet, so a handful of sections can carry a headline move.
- If a residual survives, size it with the section as the unit of analysis and report an interval that carries the design effect, roughly 1 + (m - 1) * rho for near-equal sections, rather than a learner-level interval that will overstate certainty.
Follow-up
- If the drop survives every check, how would you design the follow-up study to establish cause, and at what unit would you randomise?
- The delayed retention check is itself served by the adaptive engine. What does that do to the verified mastery definition, and how would you fix it?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
Give an example of a time you disagreed with a team member’s technical…
Give an example of a time you disagreed with a team member’s technical approach and how you reached a resolution.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Tell an executive the usage report they publish is wrong
A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.
Approach
- Size the error before raising it. Recompute the metric both ways per org: 'the definition is wrong' and 'the number is overstated by 18 points on the largest account' get very different responses.
- Keep the two defects separate, because they push in opposite directions. Roster artefacts inflate the numerator while age-gated truncation shrinks the denominator, so the net error varies by org and a single average hides that.
- Go to the report owner first with the recomputation and a proposed replacement, not to their audience. The person who published it needs the chance to correct it themselves.
- Propose a definition that is defensible and computable today: distinct learners with at least one scored submission in the quarter, with enrollment_source in ('roster_sync','admin_bulk') excluded from activation cuts, and the age_gated share stated as a footnote rather than silently dropped.
- Say what should happen to the two reports already sent, with a recommendation and an owner. The restatement is the uncomfortable part, and leaving it for someone else to raise means the work is unfinished.
Follow-up
- The executive asks you to change the definition quietly from next quarter with no restatement. What do you do?
- How would you stop this class of defect reaching a published report again?
Disagree with a product manager using evidence, not volume
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
Approach
- Make the objection falsifiable before raising it. The claim implies two testable predictions: mean calibrated difficulty of served items rose with learner ability, and accuracy is flat within ability strata.
- Compute weekly mean irt_b of served items per arm, restricted to items with calibration_n above your floor, and plot it against the accuracy series. If served difficulty tracked ability, the accuracy line carries no learning signal and you can show that rather than assert it.
- Build the metric that survives adaptivity: a small fixed-form set with (content_item_id, version_no) held constant, served to both arms, reported as the pilot's accuracy readout.
- Bring the replacement to the meeting, not only the refutation. A PM who has been told the number is meaningless still has a launch decision and no instrument.
- Separate the two questions out loud: whether the selector helps learners is open and testable; whether first-attempt accuracy measures it is settled, and it does not.
Follow-up
- The fixed-form set costs each learner six minutes a fortnight. How do you justify that to the same PM?
- Mean served irt_b is flat but accuracy still rose four points. What do you look at next?
- 01
Give an example of a time you disagreed with a team member’s technical approach and how you reached a resolution.
- 02
A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.
- 03
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
Is this an official The University of Pennsylvania interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at The University of Pennsylvania. Rounds and questions reflect what candidates have reported, not a process The University of Pennsylvania has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
The process can vary, but expect a timeline spanning several weeks from the initial screen to a final decision. Be prepared for a professional, structured, and consistent pace.
PracHub interview research ↗What is the best way to prepare for the technical rounds?
Focus on the fundamentals. Ensure you are proficient with SQL (specifically window functions) and can explain the logic behind statistical significance and A/B testing design.
PracHub interview research ↗What makes a candidate stand out?
Candidates who stand out are those who can link their technical work to the broader mission of the institution. Always frame your answers in the context of the user or the business impact.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22