At ProSidian, the Data Scientist role serves as a critical bridge between complex human capital challenges and actionable, data-driven strategy. You will operate within a high-impact consulting environment, supporting federal clients—such as the National Science Foundation (NSF)—to modernize their workforce planning, improve employee experiences, and enhance organizational decision-making through advanced analytics.
This role is not merely about building models; it is about delivering value-centric solutions that align with federal regulations and business objectives. Whether you are applying NLP and sentiment analysis to workforce surveys or building predictive models for HRStat and GPRA compliance, your work directly influences how federal agencies manage their most valuable asset: their people. You will be expected to function as a technical lead, translating raw data into executive-level insights that drive digital transformation and operational efficiency.
Working in a hybrid, mission-driven environment, you will encounter a high degree of autonomy and complexity. Success here requires a blend of deep technical proficiency in tools like, and a consulting mindset that prioritizes clear communication, stakeholder management, and the ability to solve ambiguous problems under the umbrella of commitment to excellence and integrity.
Initial Screening
reportedWhoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.
What to demonstrate
- Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
- Whether your reason for wanting the role points at the work itself rather than the company's reputation
- Whether your language signals the level being screened for: what you decided yourself versus what you were handed
How to prepare
- Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
- Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
- Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
Technical Assessment
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Behavioral Assessment
reportedThis round decides whether you owned a decision or watched one happen nearby. Interviewers for data roles listen for the point where the analysis stopped being a report and started changing what someone did, so build each story around that hinge: what was going to happen by default, what you found, and what happened instead. The most common weakness is a story that ends at delivery. If you can name the decision your work changed and the number that moved because of it, most follow-ups become easy.
What to demonstrate
- Whether the decision was yours to influence, or whether you are narrating a team outcome in the first person
- The counterfactual: what would have been done without your analysis, and why that default was worse
- How far your involvement ran past the handoff, and whether you checked that the change did what you predicted
How to prepare
- Pick three projects and write one sentence for each naming the decision-maker, the choice in front of them, and what they chose after seeing your work. If you cannot name a person and a choice, the story is not ready for this round.
- Reconstruct the baseline for your strongest project from the original query or dashboard rather than memory, so the before-number survives a follow-up asking where it came from.
- Prepare an honest version of a project where your recommendation was overruled, including what you did with the analysis afterwards.
Final Selection
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub editorial advice for the preparation topics above.
Computing days-to-pay or proposal cycle time over completed records only
At any snapshot date, invoices that have already been paid are disproportionately the fast ones, and proposals that already have a decision are disproportionately the quick ones. Averaging over the completed set alone biases both numbers downward, and the bias grows exactly when the business is deteriorating, because the slow cases are the ones still open. Unpaid and undecided records are right-censored; use Kaplan-Meier or a restricted mean up to a fixed horizon, and never fill paid_at with a placeholder.
Modelling win rate on proposals with a recorded outcome, using fields written after the decision
Two failures compound here. First, stage IN ('withdrawn','no_decision') is not missing at random: those are disproportionately deals that were going to be lost, so training on won-plus-lost only inflates apparent win rate and distorts the coefficients. Second, fields like engagement_id, final scope and revised pricing are populated after the outcome is known, so including them leaks the label and produces a model with excellent backtest accuracy and no forward value. Restrict features to values knowable at submitted_at.
Comparing periods without accounting for seasonality or day-of-week
Compare whole weeks against whole weeks and check whether the same swing appeared in prior cycles or prior years before attributing it to anything you changed. Weekday and weekend populations often differ enough that a Tuesday-to-Saturday comparison is meaningless.
Analysing at a different unit than the one randomised
Say out loud what was randomised (user, device, account, cluster) and make the analysis unit match, or account for the clustering with cluster-robust standard errors, the delta method, or aggregation up to the randomised unit. Randomising users and then running a test over sessions understates variance and inflates the false-positive rate.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How would you design an NLP pipeline to analyze sentiment in unstructu…
How would you design an NLP pipeline to analyze sentiment in unstructured employee survey feedback?
Approach
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
What specific metrics would you use to evaluate the performance of a p…
What specific metrics would you use to evaluate the performance of a predictive model for workforce planning?
Approach
- Check what information would not exist at prediction time, and exclude it.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Build a dense consultant-week hours spine with zero weeks
You are given time_entries (time_entry_id, consultant_id, engagement_id, charge_code, work_date as datetime64, hours, is_billable, status) and a list of consultant_ids in scope plus a start and end date. Produce a DataFrame with exactly one row per consultant per Monday-anchored week in that range, with columns consultant_id, week_start, billable_hours, total_hours. Weeks in which a consultant logged nothing must appear as zeros rather than be absent. Count only rows with status == 'approved'. Do not loop over consultants and do not use resample.
Approach
- Anchor weeks arithmetically: week_start = work_date - pd.to_timedelta(work_date.dt.weekday, unit='D'). Avoid dt.isocalendar().week, whose numbers restart each January and do not sort across a year boundary.
- Construct the spine from the inputs, not from the data: pd.date_range(start, end, freq='W-MON') crossed with the in-scope consultant_ids via pd.MultiIndex.from_product, so its size is fixed before any aggregation runs.
- Aggregate approved entries once with a single groupby on [consultant_id, week_start], producing total hours and billable hours (hours.where(is_billable, 0).sum()) in the same pass.
- Left-merge the aggregate onto the spine, fillna(0.0) on both hour columns, and keep float dtypes so later ratios do not integer-divide.
- Assert the shape before returning: len(out) == n_consultants * n_weeks, and out.total_hours.sum() equals the sum of hours over the filtered input inside the window.
Worked solution 20 min
- Filter to status == 'approved' and to work_date within [start, end], then derive week_start by subtracting the weekday offset.
- Build weeks = pd.date_range(start_monday, end, freq='W-MON') and spine = pd.MultiIndex.from_product([consultant_ids, weeks], names=['consultant_id','week_start']).to_frame(index=False).
- agg = filtered.assign(billable_hours=filtered.hours.where(filtered.is_billable, 0.0)).groupby(['consultant_id','week_start'], as_index=False)[['hours','billable_hours']].sum().
- out = spine.merge(agg, how='left', on=['consultant_id','week_start']).fillna({'hours':0.0,'billable_hours':0.0}).rename(columns={'hours':'total_hours'}).
- Sort by consultant_id then week_start and run the two assertions on row count and hour conservation.
Follow-up
- Utilisation gets computed on this spine. What breaks for a consultant hired mid-window or terminated mid-window, and where would you fix it?
- How would you extend this to consultant x engagement x week without the spine exploding to a row count nobody can hold in memory?
Longest consecutive bench spell per consultant from weekly activity
You have a dense series (consultant_id, week_start, billable_hours, full_week_leave BOOLEAN) covering 26 weeks for every billable-role consultant, with zero-hour weeks present as 0. Return every maximal run of consecutive weeks where billable_hours = 0, as consultant_id, spell_start_week, spell_end_week and week_count, then each consultant's longest run. Weeks flagged full_week_leave must neither break a run nor count toward its length. Mark runs still open at the last week of the series rather than reporting them as finished.
Approach
- Remove full_week_leave weeks from the sequence first, then re-number what remains with ROW_NUMBER() OVER (PARTITION BY consultant_id ORDER BY week_start). Deleting before numbering is precisely what makes a leave week transparent to a run instead of a break in it; filtering afterwards would leave a gap in the sequence and split the spell.
- Flag each surviving week as bench (billable_hours = 0) or active, then compute two row numbers per consultant: one over all surviving weeks, one partitioned additionally by the bench flag. Their difference is constant inside any run of adjacent same-flag weeks, which is the island key.
- Group the bench weeks by (consultant_id, island key) and take MIN(week_start), MAX(week_start) and COUNT(*). Do not island on week_start minus an interval times a row number: that variant assumes the spine has no holes, and the leave removal has just deliberately put holes in it.
- Pick the longest spell per consultant with ROW_NUMBER() OVER (PARTITION BY consultant_id ORDER BY week_count DESC, spell_start_week) so ties resolve deterministically, rather than a MAX that cannot carry the accompanying dates.
- Handle both edges as censoring. A spell touching the last week of the series is still open and its length is a lower bound, so flag it; a spell touching the first week may have started earlier, so either flag it too or exclude it from any length distribution and say which.
Follow-up
- A spell begins before your 26-week window. How do you detect that, and what do you report as its length?
- Rebuild this at daily grain. What changes about weekends, holidays and the island key?
- How would you separate true bench from unbilled delivery work on a live engagement, given only charge codes?
Rebuild proposal stage timeline from a mutable event stream
fct_proposal holds proposal_id, client_id, expected_value_usd, stage, is_competitive, created_at and decided_at, and its stage column is overwritten in place. fct_proposal_stage_event holds stage_event_id, proposal_id, from_stage, to_stage and occurred_at. For proposals created in the trailing year, rebuild the ordered stage timeline with days spent in each stage, flag durations still open at the snapshot as censored, and build a funnel of distinct proposals and expected value ever reaching qualifying, scoping, submitted and decided. Reconcile each timeline's terminal stage against fct_proposal.stage and report disagreements.
Approach
- Order events per proposal and derive each stage's exit with LEAD(occurred_at) OVER (PARTITION BY proposal_id ORDER BY occurred_at, stage_event_id). The tiebreak on the event key matters because two transitions can share a second. LEAD returns NULL on the last event of every proposal, decided or not, so a NULL lead marks the end of the stream and nothing more — it is not by itself a censoring indicator.
- Close that final interval by state rather than by the NULL. Where decided_at IS NOT NULL, close it at decided_at with is_censored = FALSE; where decided_at IS NULL, close it at snapshot_date with is_censored = TRUE. Only the second group is censored, and marking it lets a consumer choose Kaplan-Meier or a restricted mean. A median computed over completed transitions alone is biased downward, because at any snapshot the still-open records are disproportionately the slow ones, and the bias grows exactly when the pipeline is slowing. A decided_at earlier than the last event's occurred_at yields a negative final interval: report those proposals, do not clamp them.
- Handle re-entry. A proposal can go scoping to submitted and back to scoping, so per-stage entries outnumber proposals. Build the funnel on COUNT(DISTINCT proposal_id) for 'ever reached stage X', and keep the per-entry rows separate for duration work.
- Build the funnel from the reconstructed timeline, never from fct_proposal.stage. That column records only where a proposal came to rest, so every proposal that passed through scoping and moved on is invisible in it and the funnel's middle collapses.
- Reconcile the last to_stage per proposal against fct_proposal.stage. Disagreements mean events are missing or a row was edited outside the event path, and the size of that set is the ceiling on how far anything else here can be trusted.
- If asked for win rate by value, exclude stage IN ('withdrawn','no_decision') from both numerator and denominator and filter is_competitive = TRUE. Those outcomes are not losses, they are not missing at random, and sole-sourced follow-on work inflates the rate.
Worked solution 45 min
- Count events per proposal and check for proposals with zero events; those exist only in the mutable table and must be reported, not dropped.
- Add the LEAD-based exit timestamp with the event-key tiebreak, then close each final interval with COALESCE(lead_occurred_at, decided_at, snapshot_date) and set is_censored only where decided_at IS NULL, so days_in_stage is defined on every row.
- Aggregate to the funnel with COUNT(DISTINCT proposal_id) and SUM of expected_value_usd per 'ever reached' stage.
- Compute the median days in stage over completed transitions only, and report the censored count beside it so the reader sees what the median excludes.
- Take the last to_stage per proposal with a window function and diff it against fct_proposal.stage, returning the mismatched proposal_ids.
Follow-up
- Estimate median proposal cycle time with the censored proposals included. What does the estimator need from your output?
- Two events for one proposal share occurred_at to the second. How does your ordering resolve it, and how would you detect that it happened?
- Which fields on a won proposal are populated only after the decision, and why does that rule them out of a win-rate model?
How do you manage competing priorities when working on multiple consul…
How do you manage competing priorities when working on multiple consulting engagements?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
What does "Client Service" mean to you in the context of a long-term g…
What does "Client Service" mean to you in the context of a long-term government engagement?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you ensure your models are compliant with federal data governan…
How do you ensure your models are compliant with federal data governance and privacy standards?
Approach
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Explain the difference between supervised and unsupervised learning in…
Explain the difference between supervised and unsupervised learning in the context of workforce attrition modeling.
Approach
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Estimate a mandatory pricing review at a value threshold
Proposals with expected_value_usd at or above 250,000 must pass a mandatory pricing review before submission; below that the originating partner prices unilaterally. The rule has stood for three years, nothing was randomised, and leadership wants to know whether the review improves realisation on the resulting engagements. About 900 proposals a year land in fct_proposal, and realisation is only observable for those that reach stage = won. Design the identification strategy, state the estimand precisely, and name the two checks that would make you abandon the design.
Approach
- Set up the design. Running variable is expected_value_usd as recorded when the rule is applied, cutoff 250,000, outcome realisation on the resulting engagement. Estimate with local linear regression either side, triangular kernel, MSE-optimal bandwidth and robust bias-corrected confidence intervals. Do not fit a global high-order polynomial, which drags weight from proposals far from the cutoff onto the estimate at it.
- Check compliance before calling it sharp. If some sub-threshold proposals are reviewed voluntarily or some above-threshold ones skip review, the design is fuzzy, which is instrumental variables with the threshold indicator as the instrument. The estimate is a Wald ratio, the jump in realisation divided by the jump in review probability, and it identifies the effect for compliers at the cutoff only.
- Run the manipulation test and interpret it carefully. A density test on expected_value_usd will very likely reject, because proposal values heap at round numbers and 250,000 is exactly such a number. Heaping alone is not sorting, but a partner pricing at 249,000 to dodge review is, and both look the same in the density. Separate them by checking whether excess mass appears only just below the cutoff or at round numbers throughout the range, and by running a donut specification that drops a window around the threshold.
- Check covariate continuity at the cutoff for industry_segment, account_tier, is_competitive, practice_area and originating_partner_id, using the same local linear specification with each covariate as the outcome. A jump in any of them at 250,000 is a jump in the population rather than in the treatment, and it kills the design.
- Confront the outcome-side selection. Realisation exists only for won proposals, and review may itself move the win rate, so conditioning on a win conditions on a post-treatment variable. Estimate the RD on win rate first; if it jumps, switch to an outcome defined over all proposals, such as realised fees per proposal with losses entering as zero, or report Lee bounds rather than a point estimate.
- Size the design before committing to it. Only proposals inside the bandwidth contribute, so 2,700 proposals over three years with a bandwidth capturing perhaps 15% leaves an effective sample near 400 split across the cutoff. Compute the MDE on that number, not on 2,700, and say in advance whether it can detect the effect that would change the policy.
Worked solution 45 min
- Histogram expected_value_usd in fine bins around 250,000 and across the full range, to separate generic round-number heaping from cutoff-specific sorting.
- Estimate the first stage: probability of a recorded review as a function of the running variable, and read the jump at the cutoff. A jump well below one means fuzzy, so move to the Wald ratio.
- Run the local linear RD on realisation with an MSE-optimal bandwidth and robust bias-corrected intervals, then repeat at half and twice that bandwidth.
- Repeat the same specification with each pre-treatment covariate as the outcome, and with placebo cutoffs at 200,000 and 300,000.
- Run the RD on win rate, then on realised fees per proposal defined over all proposals, and compare the three results.
- Report the MDE computed on the in-bandwidth proposal count alongside the estimate.
Follow-up
- If the density test rejects and the donut specification moves the estimate materially, what do you conclude and what do you put in the deck?
- The estimand is local to 250,000 and leadership wants to know whether to lower the threshold to 100,000. What can you honestly say?
- The review is triggered on the value at submission, but the value is revised afterwards. Which version of the running variable do you use, and why?
Firm margin rose while every pricing model's margin fell
Firm-wide engagement gross margin rose from 31% to 34% quarter over quarter. Split by fct_engagement.pricing_model, margin fell inside time_and_materials, fixed_fee, retainer and outcome_based. Sources are fct_invoice_line (engagement_id, line_type, amount_usd, period_start_date, period_end_date) and fct_time_entry (engagement_id, hours, cost_rate_usd, status). Explain the aggregate move, quantify how much of the 3-point change is mix and how much is within-model, and state which number you put in front of practice leadership and why.
Approach
- Write the aggregate as a weighted mean: M = sum over pricing models p of w_p * m_p, where m_p is margin within p and w_p is p's share of the fee base. The weights must be the fee base, because margin is a fees-weighted ratio; weighting by engagement count answers a different question.
- Apply the exact three-term decomposition: dM = sum (w_p1 - w_p0) * m_p0 + sum w_p0 * (m_p1 - m_p0) + sum (dw_p)(dm_p). Verify the three components sum to the reported +3.0 points before interpreting any of them.
- Identify which weights moved and in which direction; the aggregate can only rise this way if fee share shifted toward the structurally higher-margin models, typically fixed_fee delivered under budget or retainer.
- Guard against composition inside each stratum: repeat the same decomposition within pricing_model by practice_area and account_tier, so a within-model decline is not itself an unexamined mix effect.
- Confirm both sides of the ratio are on the same clock and the same scope: numerator includes line_type IN ('fees','milestone','credit_note') with credits signed negative; the cost side includes all approved delivery hours, billable and non-billable alike.
- Report within-model margin as the headline delivery signal and mix as a separate, named line with its own owner, since mix is a sales decision and margin is a delivery one.
Follow-up
- Next quarter the mix reverts. What happens to the aggregate, and what will you have already told leadership so this is not a surprise?
- Which of the four pricing models should not be pooled with the others even for a mix-adjusted number, and why?
- How would you present this if the mix shift were deliberate strategy rather than accident?
Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and set a baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
- Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
- Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.
Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Worked solution ↗02One query pattern, written three times
- Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
- On the third attempt, write the grain of every CTE as a comment before writing its body.
- Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.
Deliverable: Three independent versions of the same query plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗03Only the statistics you will be asked to defend
- Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
- Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
- Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.
Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.
Practice prompt ↗Practice prompt ↗04One case, and the assumptions holding it up
- Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
- Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
- Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.
Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
- Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in a longer weekend block
- Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
- Re-read only your own notes from this week, and open no new material.
- Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
Describe a time you had to explain a complex analytical finding to a s…
Describe a time you had to explain a complex analytical finding to a stakeholder who lacked a technical background.
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Pick a story where you drove the decision, not one where you observed it.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Describe a time you utilized SQL or Python to clean and prepare a comp…
Describe a time you utilized SQL or Python to clean and prepare a complex dataset for executive reporting.
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Quantify the outcome, including what you would not claim credit for.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Defending your own impact claim without randomisation or clean units
At your review you plan to claim that the realisation dashboard you built recovered 1.4 million dollars. The evidence is that engagement-month realisation rose four points over two quarters among engagements whose leads used it. Adoption was voluntary. There are about sixty client accounts and the top five carry most fees. A new rate card shipped in the same quarter. Write the claim you can defend, the estimate you would actually produce, and what you say when asked for a causal number you cannot get.
Approach
- Name the probe: whether you can separate the number you want from the number the data supports, under review pressure, without either inflating it or retreating to saying nothing can be known.
- State both identification problems concretely. Voluntary adoption means adopting leads are plausibly the ones who already manage realisation, so the comparison is confounded at the person level. The rate card changes bill_rate_usd, which sits in the realisation denominator, so part of the four-point move is arithmetic rather than behavioural.
- Neutralise what you can. Recompute realisation with bill rates snapshotted on work_date, or hold the denominator at the old rate card, so the rate-card change cannot move the metric by construction. Then rerun the comparison.
- Get the inference right for the unit count. Cluster at client_id, not engagement, because engagements in one account share a partner, a rate card and a team. With sixty accounts and five carrying most fees, the effective cluster count is far below sixty, so report a wild cluster bootstrap interval rather than plain cluster-robust standard errors, which are biased downward in that regime.
- Report both weightings and explain the divergence: an account-weighted estimate describes the typical account, a value-weighted one describes the revenue, and if they disagree a small number of accounts is carrying the result. Then give the decision-relevant sentence: the defensible range, whether its lower bound still clears the build cost, and what a proper staggered rollout would have bought.
Follow-up
- The pre-period trends for adopters and non-adopters are not parallel. What do you report then?
- You get to design the next rollout. What do you change so the same question is answerable, without randomising individual accounts?
- 01
Describe a time you had to explain a complex analytical finding to a stakeholder who lacked a technical background.
- 02
Describe a time you utilized SQL or Python to clean and prepare a complex dataset for executive reporting.
- 03
At your review you plan to claim that the realisation dashboard you built recovered 1.4 million dollars. The evidence is that engagement-month realisation rose four points over two quarters among engagements whose leads used it. Adoption was voluntary. There are about sixty client accounts and the top five carry most fees. A new rate card shipped in the same quarter. Write the claim you can defend, the estimate you would actually produce, and what you say when asked for a causal number you cannot get.
Is this an official ProSidian interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at ProSidian. Rounds and questions reflect what candidates have reported, not a process ProSidian has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the hiring process typically take?
While timelines vary by project urgency, most candidates move through the stages within a few weeks. Maintain regular communication with your recruiter to stay updated on the status of your application.
PracHub interview research ↗Is there a specific emphasis on a particular toolset?
Yes, Python is non-negotiable for modeling. Experience with Tableau or OAS is highly valued given the focus on dashboarding and executive reporting for federal clients.
PracHub interview research ↗What is the culture like at ProSidian?
The culture is defined by the firm's eight global competencies, emphasizing Continuous Learning, Leadership, and Client Service. You will be expected to be intellectually curious, humble, and willing to question the status quo to solve complex problems.
PracHub interview research ↗How much of the work is client-facing?
As a consultant, you should expect significant client interaction. You will be the technical expert in the room, responsible for explaining how your models and insights improve the client's operational efficiency.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22