MBC · Data Scientist
Updated · 2026-09-24

MBC Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

The Data Scientist role at MBC is a pivotal position focused on delivering high-impact analytical support for critical government programs, specifically within the U.S. Navy’s SEA 21 / PAE Maritime office. You will act as the bridge between complex, large-scale datasets and actionable strategic decisions. Your work directly influences the modernization and sustainment of surface ships, requiring you to translate raw technical data into clear, persuasive briefings for government leadership.

Allocate prep to your weakest link rather than your favourite topic. Of the three things that usually gate the outcome (SQL that is correct under messy joins, sound reasoning about experiments, and structured framing of an open-ended problem), candidates tend to over-invest in modelling theory and under-invest in framing.

MBC candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Estimate eligible populations outside administrative dataDesign outputs that survive public disclosureDefend small-area rates against denominator noise

33 min read

Practice 16 Data Scientist prompts
16Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

The Data Scientist role at MBC is a pivotal position focused on delivering high-impact analytical support for critical government programs, specifically within the U.S. Navy’s SEA 21 / PAE Maritime office. You will act as the bridge between complex, large-scale datasets and actionable strategic decisions. Your work directly influences the modernization and sustainment of surface ships, requiring you to translate raw technical data into clear, persuasive briefings for government leadership.

This role is uniquely challenging because it blends advanced technical execution with significant stakeholder management and operational coordination. Beyond building machine learning models or data pipelines, you are expected to handle data calls, facilitate project meetings, and refine operational processes. Success at MBC requires a blend of rigorous analytical capability and the interpersonal finesse necessary to navigate a high-stakes, mission-driven environment.

01

Initial Screen

reported

Whoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.

What to demonstrate

  • Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
  • Whether your reason for wanting the role points at the work itself rather than the company's reputation
  • Whether your language signals the level being screened for: what you decided yourself versus what you were handed

How to prepare

  • Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
  • Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
  • Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
PracHub interview research ↗
02

Technical Deep-Dives

reported

A handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.

What to demonstrate

  • Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
  • Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
  • Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.

How to prepare

  • Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
  • Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
  • If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
PracHub interview research ↗
03

Behavioral Interviews

reported

Rounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.

What to demonstrate

  • Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
  • What you changed afterwards, stated as a check you now run rather than a lesson you now believe
  • Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you

How to prepare

  • Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
  • Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
  • Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Treating event_at in the case event log as when the event actually happened

Staff backdate to the true date of an action and enter it later, and batch loads write many rows at once, so event_at and recorded_at diverge systematically and the divergence is largest at fiscal-period and reporting-deadline boundaries. An interrupted time series keyed on event_at will read a queue flush before a deadline as a level shift caused by the policy change, and a dashboard keyed on recorded_at will show activity on days when nothing happened. Reconcile both columns, model the entry lag distribution, and state which clock each metric uses.

02

Assuming record-linkage error is random noise that averages out

False non-matches concentrate among people with name changes, transliterated or hyphenated names, unstable addresses and no durable identifier, which are the same people a disparity analysis is usually about. The linked cohort is therefore systematically more stable than the population, and any disparity estimate computed on it is biased toward finding no disparity. Carry link_method and link_confidence into the analysis, test whether results hold as the confidence threshold moves, and report the match rate by subgroup as a diagnostic rather than a footnote.

03

Never asking what decision the analysis will inform

Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.

04

Optimising accuracy on a heavily imbalanced target

State the base rate first, then choose the metric from the relative cost of a false positive against a false negative: precision and recall at the operating threshold, PR-AUC, or expected cost. At a 1 percent positive rate, predicting the majority class for everyone scores 99 percent accuracy and is worthless.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

13 technical prompts3 include a worked solution

Which key performance indicators (KPIs) would you choose to measure th…

medium
machine learning and modelling

Which key performance indicators (KPIs) would you choose to measure the success of a new maintenance scheduling algorithm?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Say how the offline result would be validated online before it is trusted.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • Where could label leakage enter this setup?

Laplace release of nested counts with consistent post-processing

hard
differential privacylaplace mechanismpost-processing

counts holds exact denial counts by (district_geo_id, tract_geo_id, denial_reason_code), where tracts nest inside districts and each constituent contributes to exactly one cell. Release the tract by reason cells and the district totals under a total privacy budget epsilon, using a Laplace mechanism you write yourself. Then post-process each district so its tract cells are non-negative and sum exactly to the released district total. Finally show empirically how per-query error moves when one epsilon is split across k sequential queries.

Approach
  1. Fix the adjacency and the sensitivity before touching data. The full tract by reason histogram has L1 sensitivity 1 under add-or-remove adjacency and 2 under replace-one, because a replaced record leaves one cell and enters another. Say which you are using; it doubles the noise scale b = sensitivity / epsilon_cells.
  2. Split the budget honestly. District totals are a coarsening of the same records rather than a disjoint query, so releasing them alongside the cells composes sequentially and epsilon_cells + epsilon_totals = epsilon. Disjointness buys you something only across records, for example separate districts released by separate mechanisms.
  3. Add noise with a seeded generator: rng.laplace(0, b, size). Laplace(b) has variance 2b squared, so per-cell standard deviation is sqrt(2) * b. Compare that number against the typical cell count before deciding the release is publishable at all.
  4. Post-process per district: clip the noisy district total at 0, then project the noisy tract vector onto {x >= 0, sum x = T} by minimising squared distance. The solution is x_i = max(y_i - tau, 0) with tau found by sorting y descending and scanning for the value that makes the sum hit T. If integers are required, round with largest remainders so the total survives the rounding.
  5. Run the sweep: for k in 1 to 8, split one epsilon into k sequential queries each with b = k * sensitivity / epsilon, and plot RMSE against k. Per-query standard deviation grows linearly in k, which is the whole argument for releasing fewer and coarser queries.
  6. State why the reconciliation is free: any function of a differentially private output is differentially private with the same parameters, provided it touches no raw data again.
Follow-up
  • You could skip epsilon_totals and derive district totals by summing the noisy tract cells. Compare the variance of that against a directly measured total and say when each wins.
  • A reason code appears in only two tracts, with a true count of 3 in each. What do you publish, and does the noise alone make it safe?
  • The same table is released monthly for a year. What is the honest statement about cumulative budget, and what would you change in the design?

Reconcile obligations, ceilings and outlays across snapshot dates

easyWorked solution
snapshotsaggregationdata quality

obligations holds one row per funding action: obligation_action_id, award_id, action_type, action_at, fiscal_year, obligated_delta_cents (signed), ceiling_amount_cents, outlay_to_date_cents, program_code, snapshot_date and is_current_snapshot. The table is a stack of as-of snapshots, and outlay_to_date_cents is an award-level cumulative figure repeated on every action row inside a snapshot. Produce a per-award frame with net obligated value, ceiling in force and outlay; a program by fiscal-year obligation total; and an exception report for awards where net obligated exceeds the ceiling or outlay exceeds net obligated.

Approach
  1. Filter to is_current_snapshot and assert that exactly one snapshot_date survives. If several do, the flag is maintained per award rather than per snapshot and every total below counts some awards twice.
  2. Net obligated per award is the sum of signed obligated_delta_cents. Deobligations and terminations are negative and must be summed in; filtering to action_type = 'new_award' reports gross intent, not money committed.
  3. Ceiling is the value carried on the latest action by action_at, because a modification can raise it. Summing ceiling_amount_cents across actions is a category error: it is a maximum the award may reach, not cash.
  4. Outlay is award-level and repeated, so reduce it with first after asserting nunique() == 1 per award. Summing it multiplies the figure by the action count, and the result still looks plausible, which is what makes it dangerous.
  5. Build fiscal totals by grouping deltas on the action's fiscal year, and cross-check the stored fiscal_year against the value derived from action_at. Disagreements are an assignment bug worth listing, not rounding.
  6. Exception frame: net > ceiling, outlay > net, earliest action not of type new_award, and net < 0. Report counts and cents rather than percentages of a denominator you have not defended.
Worked solution 25 min
  1. Filter on is_current_snapshot, then assert obligations['snapshot_date'].nunique() == 1 and stop if it fails.
  2. Group by award_id with a different reducer per column: sum for obligated_delta_cents, the value at max action_at for ceiling_amount_cents, and first for outlay_to_date_cents guarded by an nunique check.
  3. Group by program_code and fiscal_year over the signed deltas for the fiscal totals, and recompute fiscal_year from action_at to list mismatched rows.
  4. Build the exception frame by stacking the four rule violations with a rule column, so one award can appear more than once.
  5. Return the three frames with cents intact; convert to dollars only at presentation.
EXPECTED RESULTA per-award frame (award_id, net_obligated_cents, ceiling_in_force_cents, outlay_cents, outlay_ratio), a program by fiscal-year total built only from signed deltas, and an exception frame with one row per violated rule. Awards carrying terminations show net below the sum of their positive actions, and a low outlay_ratio on a recent award is expected rather than an exception.
Follow-up
  • Outlay sits at 40 percent of obligations for a program this fiscal year. Name three innocent explanations before you write the word underspending.
  • Leadership asks how much money was committed this year for an award that spans three years with option exercises. What number do you give and what do you refuse to give?
  • The snapshot flag turns out to be per award. What breaks in your aggregation and how do you rebuild it from snapshot_date alone?

For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Design one test end to end on paper
  • Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
  • Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
  • State in advance what you will do if the primary metric is flat while a secondary metric is significant.

Deliverable: A one-page test design with a decision rule written before launch.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Power arithmetic until it is automatic
  • Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
  • Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
  • Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.

Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Variance and the unit-of-analysis problem
  • Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
  • Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
  • Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.

Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.

Practice prompt ↗Practice prompt ↗
04Validity threats you can actually test for
  • Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
  • Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
  • Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.

Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05When randomization is not available
  • Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
  • Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
  • List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.

Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.

Practice prompt ↗Practice prompt ↗
06The readout query
  • Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
  • Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
  • Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.

Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.

Practice prompt ↗Practice prompt ↗
07Present it to someone who will not read the appendix
  • Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
  • Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
  • Rewrite your opening line so the recommendation lands before any methodology.

Deliverable: A one-page readout whose first line is the recommendation.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.

Tell me about a time you had to explain a complex technical finding to…

medium
behavioural and stakeholder questions

Tell me about a time you had to explain a complex technical finding to a non-technical stakeholder.

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Close with what you would do differently, concretely.
  3. Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
  • What would you do differently if you ran that project again?
  • What did you decide not to do, and why?

Explain why a ranked district table should not be published

easy
small-area estimationuncertaintyexecutive communication

You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.

Approach
  1. Convert each published margin to a standard error first: at 90 percent confidence the standard error is moe divided by 1.645. Compute the relative standard error of the denominator for every area in the table.
  2. Show where the uncertainty lives. Treating the administrative count as fixed and independent of the survey estimate, the rate's relative standard error equals the denominator's, so an area with a 24 percent relative standard error on population has a 24 percent relative standard error on its rate no matter how clean fact_service_request is.
  3. Demonstrate the instability rather than asserting it: resample each denominator from its published standard error, recompute the ranking a few thousand times, and report how often each area actually lands in the worst twelve. Areas that appear in only a third of draws are not findings.
  4. Replace the rank with something that survives the uncertainty: aggregate to a geo_level where the relative standard error clears a threshold you state out loud, or group areas into tiers whose intervals do not overlap, and name the threshold as a choice you made rather than a standard.
  5. Give the executive one sentence she can repeat without you in the room, for example that the data supports naming a group of high-demand areas but not ordering them.
Follow-up
  • Which relative standard error threshold do you use, and why is it a judgement rather than a rule?
  • Two of the areas you aggregated sit across a boundary redraw. What breaks in the join, and how would you notice?

Allocate one analyst-week across three competing requests

medium
prioritisationcapacityexpectation setting

Three requests land on Monday and you have one analyst-week. A statutory report of decision counts by program_code is due in nine days. An operations team wants a backlog projection to size next quarter's staffing. An equity lead wants the first-response equity ratio by deprivation quartile for a briefing in six weeks. Produce the allocation, the message you send to whichever teams you defer, and name the one deliverable you will refuse to produce in the time available.

Approach
  1. Sort by who owns the date. The statutory deadline is external and non-negotiable; the other two have owners who can trade scope or timing, which makes them negotiable even though both feel urgent.
  2. Cost each request in hours including the slow parts rather than the query: the statutory extract needs a validation pass and review time, the backlog projection needs arrival rates and a real capacity ceiling from dim_unit, and the equity ratio needs request-type stratification and a boundary-vintage-correct join to dim_geography.
  3. Allocate with review time inside the estimate, not after it. Finish the statutory report early enough that a second person can check it, because a late correction on a statutory return costs more than every other item on the list combined.
  4. Scope the backlog projection down to something honest: a capacity-bounded projection that states the stability condition, since a queue only clears when the arrival rate is below the throughput ceiling, and Little's Law relates work in progress to arrival rate times time in system only in steady state, which a seasonal arrival pattern violates.
  5. Turn the deferral into a commitment: a date, a named smaller interim artefact, and the specific input you need from them in the meantime, so "deferred" is falsifiable rather than a soft no.
  6. Refuse the one thing that cannot be done well: a single-date backlog clearance forecast with no capacity assumption, and say why in one sentence rather than negotiating it down.
Follow-up
  • The operations director escalates to your manager. What do you send, and what do you not say?
  • On day six the statutory extract fails a validation check. What drops, and who finds out first?
  • 01

    Tell me about a time you had to explain a complex technical finding to a non-technical stakeholder.

  • 02

    You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.

  • 03

    Three requests land on Monday and you have one analyst-week. A statutory report of decision counts by program_code is due in nine days. An operations team wants a backlog projection to size next quarter's staffing. An equity lead wants the first-response equity ratio by deprivation quartile for a briefing in six weeks. Produce the allocation, the message you send to whichever teams you defer, and name the one deliverable you will refuse to produce in the time available.

PracHub interview preparation framework ↗
Is this an official MBC interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at MBC. Rounds and questions reflect what candidates have reported, not a process MBC has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How much preparation time is typical for this role?

Most successful candidates spend 2–4 weeks reviewing technical fundamentals and practicing behavioral stories. Given the specific domain expertise required, ensure you are comfortable explaining your past projects in the context of government or large-scale operational environments.

PracHub interview research ↗
What differentiates successful candidates?

The strongest candidates balance technical depth with "client-readiness": they can explain complex statistical concepts to non-technical partners clearly and professionally.

PracHub interview research ↗
What is the culture like at MBC?

MBC describes its culture as mission-driven and valuing hard work, while also emphasizing laughter and a positive environment. The company looks for people who are proactive, willing to challenge themselves, and interested in helping its clients succeed.

PracHub interview research ↗
Are there travel requirements?

Yes, this role requires the ability and willingness to travel domestically and internationally to provide in-person support at MBC and client sites.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.