A Data Scientist at Govini occupies a critical seat at the intersection of national security, supply chain logistics, and advanced machine learning. Govini is a decision science company that provides the defense acquisition community and military leadership with data-driven insights to optimize supply chains, assess technological capabilities, and secure national security pipelines. As a Data Scientist, your primary mission is to transform massive, fragmented, and often unstructured public and proprietary datasets into highly structured, actionable intelligence.
The impact of this role is direct and far-reaching. By developing sophisticated machine learning models, entity resolution pipelines, and knowledge retrieval systems, you enable defense leaders to make strategic decisions that safeguard national interests. Whether you are optimizing complex simulations or refining search and retrieval architectures, your work directly influences the software and data assets that support critical defense programs.
This position requires a unique blend of technical mastery and domain curiosity. You will not simply run off-the-shelf models; you will design custom data processing workflows, build scalable algorithms, and solve highly ambiguous data integration challenges. It is a rigorous, high-stakes environment where analytical precision and mission alignment are equally valued.
HR Screening Call
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Deep-Dive
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Take-Home Coding Assessment
reportedBefore any modelling, the dataset is itself the first test. Take-home data usually carries something broken: rows duplicated at an unexpected grain, a join that silently drops part of the population, timestamps stored in more than one timezone, or missingness correlated with the outcome. An hour spent profiling row counts, key uniqueness and date ranges is not overhead, because it decides whether every number after it is real. What separates submissions is whether you report the defects you found and adapt the analysis to them, rather than modelling over them quietly and hoping the aggregate absorbs it.
What to demonstrate
- Whether you established the grain of each table and checked row counts after every join, and said so in the writeup
- Whether data defects you found are surfaced with their effect on the conclusion, instead of being dropped without comment
- Whether filters and exclusions are reproducible from the submitted code, with the size of the excluded population quantified
How to prepare
- Write a short profiling script you can point at any unfamiliar table: row count, distinct key count, null rate per column, and the min and max of every date field, then run it before anything else
- Write the funnel or the join chain as one query and check the row count at each grain, so a silent fan-out shows up as a number rather than as a wrong answer later
- On a past dataset, list every exclusion you applied and how many rows each one removed, then draft the single sentence about it you would put in a report
Onsite Interview Loop
reportedWhere a loop includes a partner from outside the data team, that conversation usually carries the same weight as the technical ones and gets the least preparation. The person opposite you will not follow a derivation and does not need to. They are working out whether having you involved would make their decisions better or slower. The failure mode is not being too technical. It is answering a question about a decision with a description of your method, leaving the translation to them. What they carry into the debrief is the sentence you handed them, not the analysis underneath it.
What to demonstrate
- Whether a statistical result arrives as something the partner could act on, with the one caveat that would change their decision kept and the rest left out
- Whether you can state what you need from their side, in their terms: instrumentation that does not exist yet, a definition they own, or a holdout they have to agree to
- Whether uncertainty is given as a range someone can plan against, rather than as hedging that invites them to ignore the result
- Whether you ask what decision is actually on the table before explaining anything
How to prepare
- Take a result you know well and write the version for someone who stops reading after one sentence, then the three-minute version, and check the short one is not the long one with the qualifications stripped out
- For a past project, list everything you asked a non-technical partner for and how you phrased it, then rewrite each ask so it names what goes unmeasured without it
- Practise saying where a result does not apply, out loud, in one sentence that a partner could repeat accurately to someone else
PracHub editorial advice for the preparation topics above.
Computing coverage or take-up with an administrative denominator, for example approved applications divided by all applications
That ratio measures throughput among people who already found the system, not delivery to the people entitled to the service. It gets better when outreach is cut, because the marginal applicant is the one most likely to be denied or to abandon, and it gets worse when a new access channel brings in harder cases. The correct denominator is a modelled eligible population from survey microdata run through the eligibility rules, and the gap between it and the applicant count is usually the finding.
Suppressing one small cell in a published table and leaving the row and column totals in place
A single suppressed cell is recoverable by subtracting the published cells from the published margin, so primary suppression without complementary suppression protects nothing. Two separately published tabulations of the same underlying data compose as well, meaning two individually safe releases can jointly identify a cell. Either apply complementary suppression across the whole table and check it against the margins, or use a formal mechanism with an accounted budget, remembering that Laplace noise for a count query is scaled to sensitivity over epsilon, that sequential releases add their epsilons, and that noisy counts are not automatically non-negative or additively consistent across aggregation levels until post-processed.
Accepting a metric definition without asking about the denominator
Pin down the denominator, the eligibility filter and the time window before computing anything: conversion rate per session, per user, per eligible user and per new user are four different numbers with different behaviour. Restate the definition in one sentence and get agreement before you analyse.
Ignoring interference between units in a marketplace experiment
Ask whether one unit's treatment can change another unit's outcome through shared inventory, a matching pool, a social graph or a common budget. Where it can, randomise at a level that contains the spillover, such as region or time slice, and say explicitly what that costs you in statistical power.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Point-in-time join across a type 2 dimension and boundary vintages
Three frames: applications (application_id, constituent_id, submitted_at); constituent_scd2 (constituent_id, effective_from, effective_to nullable, residence_geo_id, preferred_language_code, link_confidence); geography (geo_id, geo_code, geo_level, vintage_year, effective_from, effective_to, deprivation_quartile). geo_id is unique to a geo_code and vintage pair. Attach to each application the constituent row in force at submitted_at and the geography row whose vintage window contains submitted_at. Do not use pandas.merge_asof. Return the enriched frame, unmatched counts broken out by reason, and the count of applications whose quartile differs from a naive join on the current vintage.
Approach
- Do the interval join as a key merge plus a filter: merge on constituent_id, then keep rows where effective_from <= submitted_at and (effective_to is NaT or submitted_at < effective_to). Check the intermediate row count first, because a constituent with many versions fans out before the filter runs.
- Assert one surviving match per application. Two means overlapping effective ranges in the dimension, and taking the first hides a source bug that resurfaces later as duplicated facts in every downstream aggregate.
- Re-vintage the geography rather than trusting the foreign key. The matched residence_geo_id names a geo_code plus a vintage, so read off its geo_code and run a second interval join on geo_code against geography's own validity window. A single join on geo_id skips this step entirely.
- Keep both joins as left joins, then classify the unmatched: constituent absent from the dimension, submitted before the earliest effective_from, or geo_code not present in the vintage covering that date. Each has a different owner and a different fix.
- Fix one interval convention, half open [from, to), and use it in both joins so a record dated exactly on a boundary matches exactly one row rather than zero or two.
- Quantify the alternative: join the current-vintage geography instead and count applications whose deprivation_quartile changes. That count is the number of historical records a naive join would reassign.
Follow-up
- merge_asof with direction='backward' runs on this data and returns a row for nearly everything. What does it get wrong, and on which records?
- The same job has to run over 40 million applications and the key merge exhausts memory. What is your plan, and which correctness property must survive it?
- link_confidence is below 0.9 for 8 percent of matched constituents. How does that change what you are willing to publish from this frame, particularly a subgroup comparison?
Laplace release of nested counts with consistent post-processing
counts holds exact denial counts by (district_geo_id, tract_geo_id, denial_reason_code), where tracts nest inside districts and each constituent contributes to exactly one cell. Release the tract by reason cells and the district totals under a total privacy budget epsilon, using a Laplace mechanism you write yourself. Then post-process each district so its tract cells are non-negative and sum exactly to the released district total. Finally show empirically how per-query error moves when one epsilon is split across k sequential queries.
Approach
- Fix the adjacency and the sensitivity before touching data. The full tract by reason histogram has L1 sensitivity 1 under add-or-remove adjacency and 2 under replace-one, because a replaced record leaves one cell and enters another. Say which you are using; it doubles the noise scale b = sensitivity / epsilon_cells.
- Split the budget honestly. District totals are a coarsening of the same records rather than a disjoint query, so releasing them alongside the cells composes sequentially and epsilon_cells + epsilon_totals = epsilon. Disjointness buys you something only across records, for example separate districts released by separate mechanisms.
- Add noise with a seeded generator: rng.laplace(0, b, size). Laplace(b) has variance 2b squared, so per-cell standard deviation is sqrt(2) * b. Compare that number against the typical cell count before deciding the release is publishable at all.
- Post-process per district: clip the noisy district total at 0, then project the noisy tract vector onto {x >= 0, sum x = T} by minimising squared distance. The solution is x_i = max(y_i - tau, 0) with tau found by sorting y descending and scanning for the value that makes the sum hit T. If integers are required, round with largest remainders so the total survives the rounding.
- Run the sweep: for k in 1 to 8, split one epsilon into k sequential queries each with b = k * sensitivity / epsilon, and plot RMSE against k. Per-query standard deviation grows linearly in k, which is the whole argument for releasing fewer and coarser queries.
- State why the reconciliation is free: any function of a differentially private output is differentially private with the same parameters, provided it touches no raw data again.
Follow-up
- You could skip epsilon_totals and derive district totals by summing the noisy tract cells. Compare the variance of that against a directly measured total and say when each wins.
- A reason code appears in only two tracts, with a true count of 3 in each. What do you publish, and does the noise alone make it safe?
- The same table is released monthly for a year. What is the honest statement about cumulative budget, and what would you change in the design?
Propagate survey margins of error into small-area take-up rates
counts gives geo_id and approved_constituents from administrative records, treated as exact. geos gives geo_id, population_estimate and population_estimate_moe, published at 90 percent confidence, for tracts nested inside districts. Take-up rate is approved divided by population_estimate. Propagate the denominator uncertainty by simulation: derive the standard error from the margin of error, draw 10,000 denominators per tract, and report a 90 percent interval per tract. Then estimate the probability that the tract ranked first on point estimate is genuinely among the true top five, and name the aggregation level at which the rate's relative standard error falls below 15 percent.
Approach
- Convert the published margin to a standard error with the matching critical value: at 90 percent confidence se = moe / 1.645. Reaching for 1.96 out of habit understates the spread by about 16 percent on every tract at once.
- Simulate the denominator rather than the rate. Draw N_b per tract for B = 10,000 and form rate_b = approved / N_b. The ratio is right skewed when the denominator is noisy, so take percentile intervals instead of a symmetric point plus or minus.
- Handle non-positive draws explicitly: either truncate at a floor and report the share of draws truncated, or draw lognormal matched on the same mean and standard error. A tract where truncation bites is a tract too small to rank at all, which is itself the finding.
- For the ranking question, rank tracts inside each draw and estimate P(the point-estimate leader lands in the true top five) as the share of draws where it does. Attach the Monte Carlo standard error sqrt(p(1-p)/B) so nobody reads the probability as exact.
- Aggregate by summing counts and combining tract margins in quadrature, which is the published convention and assumes the tract estimates are independent. State that assumption, then report the first level at which relative standard error of the rate drops under 15 percent.
Worked solution 30 min
- Compute se = moe / 1.645 and the point rate, then the denominator's relative standard error se / population_estimate.
- Draw a (n_tracts, 10000) matrix of denominators with a seeded generator, apply the non-positive rule, and divide the approved counts through it.
- Take the 5th and 95th percentiles along the draw axis for per-tract intervals.
- Argsort each column to rank tracts within draws, and compute the share of draws in which the point-estimate leader sits in the top five.
- Roll tracts to district by summing counts and adding margins in quadrature, recompute relative standard error, and report the first level under 15 percent.
Follow-up
- You treated the administrative numerator as exact. Name the conditions under which that is wrong and what it does to the interval.
- A director wants the ten worst tracts published as a list. What do you publish instead, and how do you explain the substitution in one sentence?
- The denominator should be the eligible population rather than everyone resident. What changes in this computation, and what new uncertainty enters?
Reconcile backdated events against their recording timestamps
fact_case_event stores event_at and recorded_at as TIMESTAMPTZ in UTC: event_at is when staff assert the action happened, recorded_at is when the row was written, and is_backdated marks event_at falling before the date of recorded_at. The agency's business day is defined by a local timezone name supplied as a parameter. For one quarter and event_type = 'decision', return a daily series carrying decisions counted on the local event_at date, decisions counted on the local recorded_at date, the median entry lag in hours, and a flag where the two counts differ by more than 20 percent.
Approach
- Convert before truncating: (event_at AT TIME ZONE :agency_tz)::date yields the local wall-clock date, whereas event_at::date silently uses the session timezone and pushes every evening decision onto the following UTC day.
- Build the two counts on separate day keys and FULL OUTER JOIN them on the date, or aggregate from a generated date spine, because a day can have recordings with no assertions or the reverse and an inner join would drop exactly the interesting days.
- Compute the lag as EXTRACT(EPOCH FROM (recorded_at - event_at)) / 3600.0 and take percentile_cont(0.5) WITHIN GROUP (ORDER BY lag_hours); the median is the right summary because the lag distribution has a long right tail from bulk backdating.
- Key the lag to one of the two day definitions and say which, since the median lag on the recorded_at day answers a different question from the median lag on the event_at day.
- Express the divergence flag as ABS(a - b) > 0.20 * GREATEST(a, b) with a guard against a zero denominator, and expect it to cluster at fiscal-period close rather than scatter randomly.
Follow-up
- An interrupted time series keyed on event_at shows a clean level shift on the policy effective date. What would you check before believing it?
- Which of the two clocks should the published on-time decision share use, and what would you write in the footnote?
- Half the quarter's backdated rows land in the final three days. Is that fraud, a batch job, or normal practice, and how would you tell them apart from these columns alone?
Deduplicate replayed decision rows before counting denials
Batch reloads wrote decision rows into fact_case_event more than once for the same case: identical application_id and event_at, different case_event_id and recorded_at, actor_type = 'system_batch'. Using fact_case_event (case_event_id, application_id, event_type, event_at, recorded_at, actor_type) and fact_application (application_id, status, denial_reason_code, decision_at), return denial counts by denial_reason_code for one month of decision events, plus the number of applications carrying more than one decision row. Treat the earliest recorded_at as the canonical decision for each application.
Approach
- Rank the decision events with ROW_NUMBER() OVER (PARTITION BY application_id ORDER BY recorded_at, case_event_id) and keep rn = 1, using case_event_id as a deterministic tiebreaker in case two replays share a recorded_at.
- In the same window pass, compute COUNT(*) OVER (PARTITION BY application_id) AS dup_count so the duplicate count falls out of the query that fixes the duplicates, rather than requiring a second scan.
- Keep two different quantities apart: the number of applications with dup_count > 1 (what the prompt asks for) and the number of surplus rows, SUM(dup_count - 1). A group replayed three times contributes 1 to the first and 2 to the second, so the two agree only in the special case where every replayed application was written exactly twice.
- Join the deduplicated event set to fact_application on application_id only after the dedup; joining the raw event table multiplies each denial by its replay factor, and because the factor varies by load batch the inflation is not a constant you can divide out.
- Filter to status = 'denied' before grouping by denial_reason_code, otherwise approved and withdrawn rows contribute a NULL reason bucket that reads as an unlabelled denial category.
- Report procedural and substantive reason codes as separate subtotals, since evidence_not_returned and over_income describe different failures and pooling them hides which one is moving.
Worked solution 25 min
- Count rows in fact_case_event with event_type = 'decision' for the month, and COUNT(DISTINCT application_id) over the same set. Their difference is the number of surplus rows, SUM(dup_count - 1) — not the number of affected applications, which is smaller whenever any application was replayed more than twice.
- Build a CTE with ROW_NUMBER and COUNT(*) OVER the same partition, then select rn = 1.
- Aggregate the duplicate measures over the rn = 1 rows: COUNT() FILTER (WHERE dup_count > 1) is the answer the prompt wants, SUM(dup_count - 1) reconciles to step 1, and COUNT() FILTER (WHERE dup_count > 2) tells you whether the two can be quoted interchangeably.
- Join the rn = 1 set to fact_application, filter status = 'denied', group by denial_reason_code.
- Run the same aggregation against the raw event table and record how much larger the denial counts come out.
Follow-up
- Denials counted from fact_case_event exceed the count of applications with status = 'denied'. Which number do you trust, and what does the gap tell you?
- Would you dedup on earliest or latest recorded_at, and what evidence would settle it?
- The replay dropped a few decision rows entirely rather than duplicating them. How would you detect that from these two tables?
Write a Python script to parse a large, unstructured CSV file, clean m…
Write a Python script to parse a large, unstructured CSV file, clean missing values, and aggregate specific metrics without relying on heavy external frameworks.
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How would you optimize a slow-running data pipeline that processes mil…
How would you optimize a slow-running data pipeline that processes millions of rows of supply chain data in Pandas?
Approach
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Explain the memory trade-offs between using Python generators versus l…
Explain the memory trade-offs between using Python generators versus loading entire datasets into memory during a data ingestion task.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Evaluate a staggered channel rollout without a control arm
An online renewal channel went live in 9 of 46 service districts on six different dates across 14 months, ordered by district readiness. There is no holdout. Using fact_application and fact_case_event, estimate the effect on the on-time decision share, which is anchored on sla_due_at rather than decision_at. Name the estimator, state the identifying assumption and how you would probe it, and say what you would do differently if only one district had adopted.
Approach
- Refuse the default: two-way fixed effects under staggered timing with effects that change over time forms comparisons that use already-treated districts as controls, and its single coefficient is a variance-weighted average that can carry negative weights, so it is not the average effect anyone is asking for.
- Estimate group-time effects directly against a not-yet-treated comparison group, then aggregate the ATT(g,t) into an event-study path and one post-adoption summary, so every comparison is a newly treated district against a district that has not yet adopted.
- Take the assumption seriously given how adoption was ordered: readiness-ordered rollout means early districts differ from late ones in the dimension the outcome responds to, so plot pre-treatment leads and then run a sensitivity analysis over bounded violations of parallel trends, because a non-significant pre-trend test at 9 treated units proves very little.
- Fix the clock before fitting anything: define adoption on the policy go-live date, reconcile event_at against recorded_at for backdated and batch-written rows, and confirm the outcome window keys on sla_due_at so a case still undecided past its due date counts as a failure instead of leaving the sample.
- Get inference right at 46 clusters with 9 treated: cluster-robust standard errors are unreliable at that treated count, so use a wild cluster bootstrap or randomisation inference over the possible adoption schedules.
- If adoption were a single district, switch to synthetic control: fit donor weights on a long pre-period of the same outcome plus predictors, require a good pre-period fit before making any claim, and draw inference from placebo permutations across donors using the post-to-pre RMSPE ratio.
Worked solution 45 min
- Build a district by month panel of the on-time share with the denominator anchored on sla_due_at, and mark each district's adoption month g.
- Estimate ATT(g,t) using not-yet-treated districts as the comparison, then aggregate to an event-time path from -6 to +12 and to an overall post-adoption summary.
- Plot the leads, and separately report the smallest violation of parallel trends that would overturn the sign of the estimate.
- Re-run inference with a wild cluster bootstrap at the district level.
- As a robustness check, drop the two earliest adopters, which contribute the longest post window and the least comparable comparison set.
Follow-up
- The event study shows a downward lead two periods before adoption. Name three explanations and how you would tell them apart.
- Districts that adopted later had worse backlogs. Does that break parallel trends, and can any of your estimators survive it?
- How would you detect a queue flush just before go-live to clear the backlog, and what would that do to your estimate?
A district take-up rate jumps after boundaries are redrawn
A service district's take-up rate rose about 40 percent year over year while its approved-constituent count was flat. The denominator comes from dim_geography (geo_id, geo_code, geo_level, parent_geo_id, vintage_year, effective_from, effective_to, population_estimate, population_estimate_moe), joined to fact_application through dim_constituent.residence_geo_id. Boundaries were redrawn between the two years. Determine whether the rise survives a correct join, and whether any remainder is distinguishable from denominator noise at 90 percent confidence.
Approach
- Audit the join key first. geo_id is unique to the geography and vintage pair; geo_code is not, and the same code can point to a different physical area after a redraw. Any join on geo_code, or any join of current-vintage geography onto historical events, silently reassigns residents and manufactures a break.
- Re-assign every application to the vintage in force at its decision_at using effective_from and effective_to, so numerator and denominator describe the same physical area in the same year. Recompute the series and see how much of the jump survives.
- If a genuine comparison across vintages is required, build an area or housing-unit weighted crosswalk between the two vintages and state that the crosswalked figures carry an allocation assumption on top of the survey error.
- Propagate the denominator error rather than reporting the rate as exact. If the margin of error is published at 90 percent confidence, the standard error is moe divided by 1.645. Treating the administrative numerator as a complete count, the rate's standard error is the rate times the denominator's relative standard error.
- Apply a stated precision rule before ranking anything. If the denominator's relative standard error exceeds the threshold you have chosen, aggregate to parent_geo_id rather than publishing the district, and say in the note which districts were aggregated and why.
Follow-up
- A stakeholder wants the ten highest take-up districts published as a league table. What do you say, and what do you offer instead?
- How would your standard error change if the numerator were not a complete count but a linked subset with a subgroup-varying match rate?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.
How do you handle malformed rows or encoding anomalies when importing …
How do you handle malformed rows or encoding anomalies when importing heterogeneous public sector datasets?
Approach
- Close with what you would do differently, concretely.
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Correct a published duration figure you got wrong
Six weeks ago you published median days to decision as the median over applications that had a decision_at. A colleague points out that applications still in status 'submitted' or 'pending_evidence' were dropped rather than censored; the Kaplan-Meier median on the same cohort is 41 days against the 22 you published. The figure sits on a public page and was quoted in a briefing. Set out what you do in the first twenty-four hours, and draft the correction note in under 150 words.
Approach
- Quantify before you talk. Recompute with still-open applications entered as right-censored at extract_date minus submitted_at, confirm the direction and size, and check whether any decision that cited the figure would have gone differently at 41 days.
- Establish who acted on it. Tell the owner of the affected decision and the owner of the publication first, in that order, rather than broadcasting to the widest audience or waiting until you have a full remediation plan.
- Write the correction so the corrected number is the first thing a reader finds: the corrected figure with its method named, the mechanism in one clause (decided-only samples exclude the slow cases that are still open, so they are length-biased), and the exact period affected.
- Fix the pipeline rather than the cell. Put the censoring rule into the metric definition, and add a test that fails if the published duration is computed on a decided-only subset.
- Say what the incident teaches about the review step that missed it, and be specific: a reviewer checking the numerator and denominator counts would have seen the cohort shrink.
Follow-up
- If the corrected figure is worse for the agency, does any part of your process change?
- What single automated check goes into the recurring job, and what does it compare?
Explain why a ranked district table should not be published
You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.
Approach
- Convert each published margin to a standard error first: at 90 percent confidence the standard error is moe divided by 1.645. Compute the relative standard error of the denominator for every area in the table.
- Show where the uncertainty lives. Treating the administrative count as fixed and independent of the survey estimate, the rate's relative standard error equals the denominator's, so an area with a 24 percent relative standard error on population has a 24 percent relative standard error on its rate no matter how clean fact_service_request is.
- Demonstrate the instability rather than asserting it: resample each denominator from its published standard error, recompute the ranking a few thousand times, and report how often each area actually lands in the worst twelve. Areas that appear in only a third of draws are not findings.
- Replace the rank with something that survives the uncertainty: aggregate to a geo_level where the relative standard error clears a threshold you state out loud, or group areas into tiers whose intervals do not overlap, and name the threshold as a choice you made rather than a standard.
- Give the executive one sentence she can repeat without you in the room, for example that the data supports naming a group of high-demand areas but not ordering them.
Follow-up
- Which relative standard error threshold do you use, and why is it a judgement rather than a rule?
- Two of the areas you aggregated sit across a boundary redraw. What breaks in the join, and how would you notice?
- 01
How do you handle malformed rows or encoding anomalies when importing heterogeneous public sector datasets?
- 02
Six weeks ago you published median days to decision as the median over applications that had a decision_at. A colleague points out that applications still in status 'submitted' or 'pending_evidence' were dropped rather than censored; the Kaplan-Meier median on the same cohort is 41 days against the 22 you published. The figure sits on a public page and was quoted in a briefing. Set out what you do in the first twenty-four hours, and draft the correction note in under 150 words.
- 03
You have twelve small areas ranked by service requests per 1,000 residents, computed from fact_service_request counts over dim_geography.population_estimate. The estimates carry population_estimate_moe published at 90 percent confidence, and for the smallest areas the margin is close to a quarter of the estimate. An executive wants the ranked list on a slide tomorrow as "the twelve worst districts". You have five minutes with her and cannot put an equation on the slide. Say what you show instead, and what you tell her.
Is this an official Govini interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Govini. Rounds and questions reflect what candidates have reported, not a process Govini has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Govini Data Scientist interview process?
The process is generally rated as average to difficult. The primary challenges lie in the length and depth of the take-home technical assessment and the endurance required for the multi-hour, back-to-back onsite interview loop.
PracHub interview research ↗What is the focus of the take-home assessment?
The take-home task is highly practical, usually involving a Python-based CSV processing and data manipulation challenge. It is designed to evaluate your ability to write clean, modular, and efficient code when dealing with messy, real-world data structures.
PracHub interview research ↗How long does the take-home assessment typically take?
Candidates report that the assessment is comprehensive and can take anywhere from 6 to 10+ hours to complete thoroughly. It is highly recommended to manage your time carefully and prioritize code quality, documentation, and error handling.
PracHub interview research ↗What should I expect during the onsite interview loop?
The onsite loop is intensive, often lasting up to five hours. It consists of back-to-back panels covering technical coding, system architecture, project deep dives, and behavioral fit. Be prepared for a continuous schedule and maintain your energy throughout the day.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22