A Data Scientist at The College Board plays a pivotal role in shaping the landscape of modern education. As a mission-driven, non-profit organization, The College Board connects millions of students to college success and opportunity through major programs like the SAT, the PSAT, and the Advanced Placement (AP) program. In this role, you will not just analyze data; you will extract actionable insights from some of the most comprehensive educational datasets in the United States to drive equity, access, and academic excellence.
The work of a Data Scientist here directly influences educational policy, student readiness indicators, and institutional planning. You will collaborate with psychometricians, product managers, and educational researchers to design statistical models, build predictive pipelines, and evaluate the efficacy of learning tools. Your models will help predict college success, optimize testing delivery, and identify systemic barriers to higher education, making your analytical contributions highly impactful and socially meaningful.
This is an intellectually rigorous environment where advanced statistical methodologies meet real-world social impact. The scale of the data is massive, requiring robust technical capabilities alongside a deep commitment to the organization's educational mission. Candidates who thrive here are those who love solving complex, ambiguous data challenges and are motivated by the prospect of using data to create a more equitable educational system.
Initial Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Screening
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Onsite Interview Loop
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub editorial advice for the preparation topics above.
Pre/post gain studies that select on low pretest scores manufacture improvement.
Any measure with reliability below 1 produces regression to the mean, so a group chosen for scoring in the bottom quartile will score higher on retest with no intervention at all. The apparent gain scales with measurement error, which for a short quiz is large. The fix is a control group selected by the identical rule, or a design that models the pretest as a covariate rather than as a selection filter.
Roster sync creates and deactivates accounts in bulk, and it is not behaviour.
An automated roster import at term start can create thousands of accounts in one hour and deactivate thousands more at term end. Signup, activation, and churn series computed without filtering enrollment_source IN ('roster_sync','admin_bulk') will show spikes and cliffs driven entirely by an integration job. It also breaks cohort retention: a roster-created account that never activates is a provisioning artefact, not a churned learner.
Treating a non-significant result as proof of no effect
Say whether the confidence interval excludes the effect sizes you would have cared about. If it does not, the honest reading is that the test was underpowered, so report the minimum detectable effect the design could have found and what sample size would resolve it.
Accepting a metric definition without asking about the denominator
Pin down the denominator, the eligibility filter and the time window before computing anything: conversion rate per session, per user, per eligible user and per new user are four different numbers with different behaviour. Restate the definition in one sentence and get agreement before you analyse.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
What is the probability of a student guessing correctly on a five-opti…
What is the probability of a student guessing correctly on a five-option multiple-choice question three times in a row?
Approach
- Write down the assumption the method needs before you use the method.
- Say what the estimate is of, and over what population it generalises.
- Translate the result into the decision it informs, in one plain sentence.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
How would you design a statistical study to measure whether a specific…
How would you design a statistical study to measure whether a specific test question contains bias against a particular demographic group?
Approach
- Say what the estimate is of, and over what population it generalises.
- Write down the assumption the method needs before you use the method.
- Translate the result into the decision it informs, in one plain sentence.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Attach the calibration in force without merge_asof
You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.
Approach
- Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
- Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
- Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
- Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
- Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
- Take the parameters positionally with np.take and write NaN where any invalidation fired.
Worked solution 30 min
- codes = pd.factorize on the concatenated item ids, applied to both frames so the mapping is shared.
- cal_key = cal_code.astype('int64') * (1 << 32) + valid_from_seconds; resp_key built the same way; sort cal by cal_key.
- pos = np.searchsorted(cal_key, resp_key, side='right') - 1.
- valid = (pos >= 0) & (cal_code[pos] == resp_code) & (cal_valid_to_seconds[pos].isna() | (resp_seconds < cal_valid_to_seconds[pos])).
- Assign irt_a and irt_b via np.take(pos) where valid, NaN elsewhere, and report the NaN count by reason.
Follow-up
- A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
- The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
- How would you test this without a golden output to compare against?
Write a SQL query to find the department with the highest average scor…
Write a SQL query to find the department with the highest average score on a specific standardized assessment.
Approach
- Say which table is the grain you start from, and join outward from it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
How would you use a SQL query to identify students who have taken mult…
How would you use a SQL query to identify students who have taken multiple advanced placement exams and calculate their average score?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Write a query to identify duplicate student registration records based…
Write a query to identify duplicate student registration records based on name and birthdate, keeping only the earliest registration.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Explain the difference between a left join and an inner join, and desc…
Explain the difference between a left join and an inner join, and describe a scenario in education data where using the wrong join would bias your analysis.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Consecutive instructional-week streaks and term-over-term return
Using fct_lesson_activity(learner_id, enrollment_id, activity_date, active_seconds) and a calendar dim_term_week(term_id, term_week_no, week_start_date, week_end_date, is_instructional), where holiday weeks carry is_instructional = FALSE, compute for each learner and term the longest run of consecutive instructional weeks containing at least one activity day. A non-instructional week must not break a run, and activity inside one neither extends nor starts a run. Then compute term-over-term return rate into term T+1 with a floor of five activity days in term T. Return learner_id, term_id, activity_days counted over the whole term including non-instructional weeks, longest_streak_weeks and the return flag.
Approach
- Collapse activity to (learner_id, term_id, term_week_no) with COUNT(DISTINCT activity_date). fct_lesson_activity is one row per item per session, so a learner with twelve rows on one afternoon is one activity day and counting rows inflates everything downstream.
- Re-index before differencing. Apply DENSE_RANK() OVER (PARTITION BY term_id ORDER BY term_week_no) across instructional weeks only, producing a gap-free sequence where a holiday week has simply been removed. Differencing raw term_week_no instead splits a run at every holiday, which is exactly the artefact the task forbids.
- Accept what dropping those weeks costs. Activity in a non-instructional week is removed from the streak calculation entirely, so a learner whose only activity lands in a holiday week has activity_days >= 1 and longest_streak_weeks = 0. That is what a streak over instructional weeks means, but it is why the two columns can disagree and why the checks must expect a zero streak alongside non-zero activity rather than assert a floor of 1.
- Apply the island trick on the re-indexed sequence: instructional_index - ROW_NUMBER() OVER (PARTITION BY learner_id, term_id ORDER BY instructional_index) is constant within a run. Group on that constant, count rows per group, take the max per learner and term.
- Build activity_days from the unfiltered term aggregate, then LEFT JOIN the streak result onto it and COALESCE longest_streak_weeks to 0. An inner join here deletes the holiday-only learners from the output and from the return-rate denominator, which is the same bug as the streak-floor assertion wearing different clothes.
- Set the activity floor on COUNT(DISTINCT activity_date) >= 5 across term T, applied to the return-rate denominator only. The floor exists to remove provisioned-but-unused accounts, so leaking it into the numerator silently redefines the metric as active in both terms. Note that a learner can clear the floor on holiday-week activity alone and enter the denominator with a zero streak.
- Define the return flag as EXISTS any activity row in term T+1 for that learner, pairing terms by an explicit term ordering rather than date arithmetic, since term lengths differ. Learners whose org has no term T+1 loaded must be excluded and counted, not treated as non-returners.
Worked solution 40 min
- Build the learner-week activity CTE with COUNT(DISTINCT activity_date) and confirm one row per (learner_id, term_id, term_week_no).
- Build the instructional re-index from dim_term_week and join activity onto it, dropping non-instructional weeks entirely.
- Apply the index-minus-row-number island grouping and take MAX(run_length) per learner and term.
- Compute activity_days per learner and term over the whole term, LEFT JOIN the streak onto it with COALESCE to 0, and count the learners who come out with activity_days > 0 and longest_streak_weeks = 0.
- Apply the five-day floor to the denominator and attach the term T+1 existence flag.
- Build a two-row fixture by hand and verify the holiday behaviour before trusting the full run.
Follow-up
- This streak rewards a learner doing one minute a week over one doing four hours in a single week. Which of those is the product claim, and what would you pair the streak with to catch the difference?
- Roster sync deactivates accounts in bulk at term end. How does that interact with the five-day floor and with the return denominator?
- Show how return rate moves as the floor goes from one day to five, and argue which floor is the honest one to publish.
How would you estimate the impact of a fee-waiver program on exam regi…
How would you estimate the impact of a fee-waiver program on exam registration rates when you cannot conduct a randomized controlled trial?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
Repair an analysis plan that peeks daily across fourteen metrics
A four-week section-randomised test is read every morning on a dashboard of 14 metrics, each with a 95% interval, and the team stops the test the moment any one of them turns significant. Two of the 14 are guardrails for harm. Quantify the false-positive inflation coming from the repeated looks and from the metric count separately. Then write the corrected plan: what is tested at what level, what correction each class of metric gets, and how early stopping is permitted at all.
Approach
- Separate the two inflations rather than lumping them. Repeated significance testing on accumulating normal data at a nominal 0.05 reaches roughly 0.14 type I error after five equally spaced looks and roughly 0.19 after ten (Armitage, McPherson and Rowe, 1969); 28 daily looks is worse than both.
- Compute the metric-count inflation: 14 roughly independent metrics at alpha = 0.05 give a family-wise error of 1 - 0.95^14 = 0.512. The two inflations compound, so the dashboard as operated is close to guaranteed to produce a win.
- Name one primary metric in advance and test it once at a pre-registered horizon at alpha = 0.05. Everything else is descriptive unless it gates the launch decision, and what gates the decision must be declared before data exists.
- If early stopping is genuinely needed, buy it with a group-sequential boundary at looks fixed in advance. An O'Brien-Fleming boundary spends very little alpha at early looks and leaves close to the full 0.05 at the final one, which suits a test you expect to run to horizon. If the team insists on continuous monitoring, use an always-valid confidence sequence and accept its sample-size cost.
- Correct the 11 secondary metrics with Benjamini-Hochberg at a false-discovery rate of 0.10. Do not correct the two guardrails: correction lowers sensitivity, and a harm check needs the opposite. Leave them at 0.05 or looser with a pre-declared harm threshold that stops the test.
- Put all of it in writing before the first look: horizon, primary metric, boundary, correction scheme, stopping rules.
Worked solution 25 min
- State the repeated-testing inflation from the accumulating-data result: about 0.14 at five looks, about 0.19 at ten, and note that 28 daily looks exceeds both.
- Compute the family-wise rate across metrics: 0.95^14 = 0.4877, so 1 - 0.4877 = 0.512.
- Select the single primary metric and fix the horizon at four weeks, or at its exposure-week equivalent if opt-in is staggered.
- Specify an O'Brien-Fleming boundary with three interim looks and record the nominal alpha at each look in the plan.
- Assign Benjamini-Hochberg at a false-discovery rate of 0.10 to the 11 secondaries and leave the 2 guardrails uncorrected with a written harm threshold.
Follow-up
- The team will not give up daily reads. Which always-valid method do you hand them, and what does it cost in sample size relative to a fixed-horizon test?
- Why do you refuse to Bonferroni-correct the guardrails when you correct the secondaries?
- The primary metric is null but three secondaries survive Benjamini-Hochberg. What do you tell the product lead, and what would change your answer?
Net revenue retention slid nine points in one quarter
Net revenue retention on institutional licenses fell from 104 percent to 95 percent quarter over quarter. Source: fct_subscription_period (subscription_period_id, account_id, org_id, plan_tier, billing_interval, seats_purchased, seats_provisioned, mrr_usd, period_start, period_end, is_renewal, status, canceled_at, churn_reason). Decompose the nine points into expansion, contraction and non-renewal, and establish whether renewal behaviour changed at all. Deliverable: the decomposition, the cohort definition you used, and a call on whether this is a trend or account news.
Approach
- Pin the cohort in writing before computing anything: orgs holding an active period twelve months before each renewal date, new logos excluded from both numerator and denominator, and non-renewals entering the numerator as zero rather than being dropped. Dropping non-renewals is the most common way this metric is overstated, and correcting it can move the headline by more than the effect under investigation.
- Rebuild the ratio from components on that cohort: starting annualised revenue, expansion, contraction, non-renewal. Check the components sum back to the reported ratio in both quarters; a residual means the two quarters were not computed on the same cohort rule and the comparison is void before any interpretation.
- Test cohort comparability on billing_interval. Multi-year contracts do not present a renewal in every window, so a quarter whose renewal cohort holds a different multi_year share is not comparable to its predecessor. Recompute both quarters holding the interval mix fixed and report that alongside the raw number.
- Rank per-org contributions to the nine points and state how much the top three carry. Institutional revenue is concentrated, so a decomposition that leaves the concentration unstated invites a trend reading of what is account news.
- Audit the inputs: confirm mrr_usd normalisation is consistent (annual divided by twelve) and look for overlapping period_start and period_end within a single account_id, which double counts revenue after a mid-term upgrade.
- Close the loop to something actionable by joining the lost orgs back to seat activation and pace adherence in the preceding term, and say whether the signal was visible early enough for anyone to have acted on it.
Follow-up
- Build the renewal-risk view from this: which signals are actionable while there is still time to act, and at what days-to-renewal would you surface an org?
- An org contracts from 900 seats to 600 but moves to a higher tier and raises MRR. Expansion or contraction, and does your decomposition handle it without double counting?
- How would you present NRR so that a quarter with an unusual renewal cohort is not read as a trend by an executive audience?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.
Walk me through a complex data science project you led during a previo…
Walk me through a complex data science project you led during a previous internship or academic program, focusing on the statistical decisions you made.
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Disagree with a product manager using evidence, not volume
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
Approach
- Make the objection falsifiable before raising it. The claim implies two testable predictions: mean calibrated difficulty of served items rose with learner ability, and accuracy is flat within ability strata.
- Compute weekly mean irt_b of served items per arm, restricted to items with calibration_n above your floor, and plot it against the accuracy series. If served difficulty tracked ability, the accuracy line carries no learning signal and you can show that rather than assert it.
- Build the metric that survives adaptivity: a small fixed-form set with (content_item_id, version_no) held constant, served to both arms, reported as the pilot's accuracy readout.
- Bring the replacement to the meeting, not only the refutation. A PM who has been told the number is meaningless still has a launch decision and no instrument.
- Separate the two questions out loud: whether the selector helps learners is open and testable; whether first-attempt accuracy measures it is settled, and it does not.
Follow-up
- The fixed-form set costs each learner six minutes a fortnight. How do you justify that to the same PM?
- Mean served irt_b is flat but accuracy still rose four points. What do you look at next?
State the impact of your last project without inflating it
Pick one project from the past year and account for its impact as if the listener could audit every number. Cover the metric you claim to have moved and its exact definition, how much of the movement is attributable to your work, what else was changing at the same time, what the counterfactual was and where it came from, and what you would have needed to measure at the start to make the claim clean. If the effect is not separable, say so and say what you would do differently. Probed: attribution discipline applied to your own work.
Approach
- Name the metric by its definition rather than its label. 'Verified mastery rate' means nothing without its numerator, denominator, window, and the reporting lag the delayed retention check forces.
- Separate the three claims usually merged into one: the metric moved, your work moved it, and the movement was worth what it cost.
- State the counterfactual explicitly and say what produced it: a holdout, matched sections, or a pre-period trend. If it was a before-and-after with no control, say that in those words.
- List the co-occurring changes with dates: a term boundary, a roster sync at scale, a pricing change, another team shipping into the same surface. In this domain the academic calendar alone routinely dwarfs a product effect.
- Convert the shortfall into a specific design artefact: the holdout you would have reserved, the instrumentation you would have shipped before the change, the analysis plan you would have written first.
Follow-up
- What fraction of the observed movement would you defend under cross-examination, and on what evidence?
- Your project shipped in term-week one. How does that change what you can claim?
- If the effect is genuinely not separable, was the project worth doing?
- 01
Walk me through a complex data science project you led during a previous internship or academic program, focusing on the statistical decisions you made.
- 02
A product manager wants to ship an adaptive practice selector to all grade bands on the strength of a four-point rise in first-attempt accuracy during a six-week pilot. You believe the rise is an artefact of the selector's target success rate. You have fct_assessment_response, dim_content_item with irt_a, irt_b and calibration_n, and the pilot arm assignment. Deliverable: the analysis that tests your objection, and how you present it so the PM can change position without it reading as a defeat. Probed: whether you make disagreement falsifiable rather than rhetorical.
- 03
Pick one project from the past year and account for its impact as if the listener could audit every number. Cover the metric you claim to have moved and its exact definition, how much of the movement is attributable to your work, what else was changing at the same time, what the counterfactual was and where it came from, and what you would have needed to measure at the start to make the claim clean. If the effect is not separable, say so and say what you would do differently. Probed: attribution discipline applied to your own work.
Is this an official The College Board interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at The College Board. Rounds and questions reflect what candidates have reported, not a process The College Board has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the Data Scientist interview at The College Board compared to big tech companies?
The interview is highly technical but focuses much more on classical statistics, probability, experimental design, and SQL rather than complex machine learning algorithms or heavy software engineering. You will be evaluated on your ability to apply statistical rigor to real-world educational problems.
PracHub interview research ↗What is the company culture like for data scientists?
The culture is highly collaborative, mission-driven, and intellectually curious. Data scientists work alongside researchers and educators who are deeply passionate about student success, resulting in a supportive and purpose-driven work environment.
PracHub interview research ↗How quickly does the hiring team follow up after each interview stage?
The College Board is known for running a very structured and punctual interview process. Candidates typically receive feedback or next steps within a week of completing an interview round, and the scheduling team is highly responsive.
PracHub interview research ↗Where are the Data Scientist roles located, and is there a hybrid work policy?
While The College Board has major offices in Washington, DC, and New York, NY, many of their data science and technology teams operate under flexible hybrid or fully remote arrangements depending on the specific team and location.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22