As a Data Scientist at The Voleon Group, you operate at the intersection of state-of-the-art artificial intelligence, machine learning, and complex financial systems. This role is crucial for turning raw, messy, and unstructured real-world data into high-quality predictive signals that drive multibillion-dollar investment strategies. You will work alongside internationally recognized experts in AI and experienced technologists, designing analytical systems, conducting rigorous exploratory data analysis, and ensuring data health across live production feeds.
Your day-to-day impact centers on building robust data pipelines, validating features through disciplined statistical frameworks, and investigating production anomalies before they affect performance. Whether you are partnering with research teams to refine predictive features or developing automated monitoring tools, your insights have a direct, traceable path to the firm's core decision-making architecture. The work demands deep curiosity, mathematical rigor, and the ability to extract clarity from ambiguous, noisy datasets.
The environment at is intensely analytical, collaborative, and fast-paced. You will be expected to combine financial intuition with advanced technical execution, leveraging modern tooling like Python, SQL, and Unix environments alongside AI-powered coding assistants. While the expectations are exceptionally high, successful candidates find immense satisfaction in tackling complex quantitative puzzles that few other firms operate at scale to solve.
Recruiter Discussion
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Triage Phase
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Technical Assessments
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Live Coding Exercises
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Behavioral Rounds
reportedBehavioural answers from data candidates get audited in a way that answers from other roles do not. When you say a model lifted retention, the next question is the denominator, the window, and how you knew the lift was not seasonal. So attach the measurement to each claim while you tell it: what the metric was before, over what period, and against what comparison. Numbers with no baseline read as rounded-up memory, and one unsupported figure tends to make the rest of the story sound rehearsed.
What to demonstrate
- Whether each impact number arrives with a baseline, a window and a comparison, or as a bare percentage
- Whether you can name the method that attributed the effect to your work (an experiment, a staged rollout, a seasonal control) or concede the link was correlational
- Whether the magnitudes stay internally consistent when the interviewer multiplies them against the scale you described earlier
How to prepare
- For each story, write the impact line as metric, value before, value after, window, and how attribution was established. Any line missing two of those five is a follow-up you will answer badly.
- Re-derive one headline number from the source table rather than the deck that reported it. Resume numbers drift upward across retellings.
- Decide in advance which figures you cannot share, and prepare the ratio or relative change you can give instead, so a confidentiality limit does not read as evasion.
Final Review Rounds
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
13 candidate reports. Individual accounts describe a particular role and hiring cycle.
Voleon Quantitative Researcher Interview Experience — Regression and Matrix Questions
The author’s Voleon phone-screen report combines mathematical reasoning, statistical modeling, and a short programming task. Matrix questions concerned the implications of unit row and column totals, including whether nilpotence was possible. The regression discussion examined bounds on the coefficient of determination and what could happen when two predictors were used together. Further question…
Read full experienceVoleon Machine Learning Engineer Interview Experience — A Time-Series Pipeline Coding Screen
After recruiter outreach, the author agreed to interview with Voleon and reports being rejected after a 90-minute machine-learning coding session. The exercise involved a pipeline with roughly five tasks and a CSV dataset containing time-series observations. The author remembers debugging, changing model parameters, running results, and extending the pipeline for inference, but not the remaining…
Read full experienceVoleon Data Scientist interview with Python modeling and analysis
I went through a thorough, structured interview process with Voleon. It started with an initial interview, followed by two technical interviews, two soft-skill interviews, and a final tech-focused interview. There were a lot of rounds, and I got the sense that they were trying to cover several different areas instead of testing only one skill. The technical interviews involved Python coding tasks…
Read full experienceVoleon Data Scientist live coding interview focused on pandas
My main interview experience was a live coding session focused on data manipulation in Python, especially pandas. There wasn’t much guidance, and I couldn’t remember every detail afterward, but I clearly remember that I didn’t produce something the interviewer seemed satisfied with. What threw me off was not knowing what level of polish or correctness they expected. I kept building and iterating,…
Read full experienceVoleon Data Scientist interview focused on live coding
My interview process focused heavily on live coding rather than a lot of conversational questions. The main topics were Python operations and basic statistics and probability. The way they tested me made it feel like they wanted to see how I actually manipulated data instead of hearing me discuss it abstractly. I expected straightforward technical work. The coding itself didn't feel impossible, b…
Read full experiencePracHub editorial advice for the preparation topics above.
Treating last-touch attribution as the causal value of a channel
The attribution label on dim_user is the output of a rule that assigns full credit to whichever touch happened to be recorded last inside a lookback window, and that rule systematically rewards channels that sit close to the conversion, especially branded search and retargeting, which largely intercept demand that already existed. Reallocating spend on those labels moves budget toward the channels that are best at being last, which is why attributed return on ad spend often improves while total signups do not. Nothing in the touchpoint data can settle this, because the counterfactual of not running the channel was never observed. The credible reads are a geo holdout or a scheduled pause, sized in advance on the total-signups metric rather than on the attributed one, and the honest framing in the meantime is that the label describes correlation with conversion and not incremental contribution.
Comparing cohort retention curves of different maturities, or building the curve from users who are still present
A cohort four weeks old has no week-8 value, so an average taken across cohorts silently drops young cohorts from the later columns and keeps them in the earlier ones. The curve then bends upward at the tail, and the reading that 'retention is improving over time' is an artefact of which cohorts survived to be measured. The same error appears in the denominator when retention is computed over users active in the current period rather than over the full original cohort, which conditions on survival and guarantees a flattering number. The fix is a triangle: fix the cohort at signup, bound every window on both sides, and only compare cells where every cohort has had the full elapsed time, publishing the rest as blank rather than as a partial average.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
Extrapolating a first-week lift inflated by novelty effects
Plot the treatment effect by days since first exposure instead of quoting one pooled average. A lift that decays toward zero across the test window is behaviour that will not persist, and annualising it produces a forecast that misses by an order of magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Solve a probability puzzle involving conditional expectation and symme…
Solve a probability puzzle involving conditional expectation and symmetry.
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Sanity-check the answer against a simple bound or a simulated case.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
Discuss how you approach feature selection for models operating in noi…
Discuss how you approach feature selection for models operating in noisy, low-signal-to-noise environments.
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- What would you monitor after launch to know the model is still valid?
Split a pooled conversion drop into rate and mix
You have weekly visit-to-signup counts by segment: a DataFrame with week, device_type, referrer_channel, visitors and signups. The pooled rate fell 0.84 percentage points between two consecutive weeks while several individual segments rose. Write a function that, for a caller-supplied list of segment columns, splits the pooled change into a rate effect, a mix effect and an interaction term that sum exactly to the observed change. Return those three scalars plus a per-segment contribution table sorted by absolute contribution, so the largest single driver can be named.
Approach
- State the algebra before coding: the pooled rate is r = sum over segments of w_s * r_s, with w_s the segment's share of the denominator. Then r1 - r0 decomposes exactly into sum(w_s0 * (r_s1 - r_s0)) for rate, sum((w_s1 - w_s0) * r_s0) for mix, and sum((w_s1 - w_s0) * (r_s1 - r_s0)) for interaction. The identity is per-segment, so it holds for any numbers you put in the four slots.
- Pivot both weeks onto a common segment index with an outer join so a segment that appeared or vanished is kept rather than dropped, then decide what rate to give a segment with no visitors in one of the weeks, and document the choice. The identity stays exact either way because the missing week's weight is 0, but the attribution does not. Filling the missing rate with 0 sends an appearing segment's entire w_s1 * r_s1 into the interaction term, since w_s0 = 0 makes both the rate term and the mix term (w_s1 - w_s0) * r_s0 identically zero; a vanishing segment then splits as -w_s0 * r_s0 in rate, -w_s0 * r_s0 in mix and +w_s0 * r_s0 in interaction.
- The convention used below instead imputes the missing week's rate as that week's pooled rate. A vanishing segment then lands wholly in mix at -w_s0 * r_s0, with rate and interaction cancelling; an appearing segment puts w_s1 * r_pooled0 in mix (volume arriving at the average rate) and only w_s1 * (r_s1 - r_pooled0) in interaction (its rate differing from that average). Impute by which week the segment is missing from, never by argument order, or the swap identities below stop holding.
- Guard the division where visitors is 0 so no NaN enters the vectors, because a single NaN poisons every sum. A segment with zero visitors in both weeks contributes exactly 0 and can be dropped; a segment missing from only one week does not contribute 0, and where its contribution lands is settled by the convention above, not by the guard.
- Compute the three components as vectors over segments, then sum. Keep the vectors, because the per-segment contribution table is what turns the decomposition into an explanation.
- Assert that the three components sum to the observed pooled change within floating-point tolerance. This identity is exact, so a mismatch means an implementation bug, not a modelling judgement.
Worked solution 25 min
- Aggregate to one row per (week, segment tuple) with summed visitors and signups, then split into w0 and w1 frames and align with an outer join, filling missing visitors and signups with 0.
- Compute w_s = visitors / visitors.sum() within each week, and r_s = signups / visitors only where visitors > 0. Where a week's visitors are 0, set that week's r_s to that week's pooled rate, the stated convention; never leave it NaN.
- rate_effect = (w0 * (r1 - r0)).sum(); mix_effect = ((w1 - w0) * r0).sum(); interaction = ((w1 - w0) * (r1 - r0)).sum().
- contribution = w0*(r1-r0) + (w1-w0)r0 + (w1-w0)(r1-r0) per segment, which reduces to w1r1 - w0r0; sort by abs and return the head.
- assert abs(rate + mix + interaction - (pooled1 - pooled0)) < 1e-12.
Follow-up
- The mix effect accounts for 0.71 of the 0.84 point drop, driven by paid_social volume. What is your recommendation, and what would change it?
- Why is a two-way split into a counterfactual rate and a residual also exact, and when would you prefer it to the three-way version?
- Segmenting on device and channel leaves a large interaction term. What does that tell you about the choice of segments?
How would you write a bash script combined with awk or grep to quickly…
How would you write a bash script combined with awk or grep to quickly parse error logs and summarize data ingestion failures?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Write a SQL query utilizing window functions to calculate moving avera…
Write a SQL query utilizing window functions to calculate moving averages and detect rolling outliers in a high-frequency market dataset.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Attribute account MRR to first-touch channel without seat fan-out
dim_user holds user_id, account_id, first_touch_channel, account_created_at_utc. dim_account holds account_id, lifecycle_status. fct_subscription_period holds account_id, period_start_utc, period_end_utc, period_status, mrr_cents_constant_fx. For a given as-of date, report active MRR by the first-touch channel of the account's earliest-created user. An account has many users and many periods. The channel totals must sum exactly to total active MRR on that date, and accounts whose label is NULL must appear as their own bucket.
Approach
- Reduce the label side to one row per account before any join: ROW_NUMBER() OVER (PARTITION BY account_id ORDER BY account_created_at_utc, user_id) = 1 over dim_user, carrying first_touch_channel. Joining dim_user to the period table directly repeats each revenue row once per seat.
- Reduce the money side to one row per account too: the period where period_start_utc <= :as_of AND period_end_utc > :as_of AND period_status IN ('active','past_due'). If more than one row survives that filter the account holds overlapping subscriptions, which must be deduped or deliberately summed, and either way the choice is stated rather than left to the join.
- Join the two one-row-per-account CTEs and SUM(mrr_cents_constant_fx). Use the constant-FX column, since the same query run in two quarters otherwise reports currency moves as channel performance.
- Bucket NULL first_touch_channel as 'unattributed' with COALESCE instead of letting it drop. Roughly the share of users with no observed touch is not small, and dropping it understates the total and inflates every named channel's share.
- Reconcile: the SUM over the joined result must equal the SUM over the period CTE alone. That single equality check catches fan-out immediately, which is why it is worth writing before presenting the number.
Worked solution 30 min
- Compute total active MRR on the as-of date from the period table alone and write the number down as the target.
- Build the label CTE and assert COUNT(*) = COUNT(DISTINCT account_id).
- Build the as-of period CTE and assert the same, investigating any account with more than one surviving row.
- Join, COALESCE the channel, aggregate, and compare the total to the target.
- Deliberately run the naive dim_user join and record the inflation factor, which is the MRR-weighted mean seat count.
Follow-up
- Which user should own the account's label: earliest created, the 'owner' role, or the one who converted? What does each choice systematically favour?
- What does this table entitle you to say about the value of a channel, and what does it not?
- Some accounts have a parent_account_id. How does rolling those up change both the numerator and the channel mix?
How do you prioritize your work when managing multiple competing data …
How do you prioritize your work when managing multiple competing data pipelines and research requests?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How do you define success when curating a brand-new third-party datase…
How do you define success when curating a brand-new third-party dataset for predictive modeling?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
What metrics would you track to evaluate the ongoing health and correc…
What metrics would you track to evaluate the ongoing health and correctness of an automated data ingestion pipeline?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How would you design a metric to capture the stability and stationarit…
How would you design a metric to capture the stability and stationarity of a financial feature over multiple market regimes?
Approach
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What are the most common experimentation pitfalls when dealing with no…
What are the most common experimentation pitfalls when dealing with non-stationary data streams?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
How do you establish rigorous stopping rules for an experiment without…
How do you establish rigorous stopping rules for an experiment without inflating your false positive rate?
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
If a key feature's predictive performance suddenly degrades in product…
If a key feature's predictive performance suddenly degrades in production, how would you systematically diagnose the drop and isolate root causes?
Approach
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Reminder volume where the guardrail opposes the primary metric
The growth team proposes tripling weekly reminder email volume. The north star is weekly active accounts completing a core action, counted on account_id from fct_event (is_core_action, account_id, occurred_at_utc, surface). Its stated guardrail is week-over-week repeat rate together with notification opt-out and unsubscribe rates. Reminders will move the primary up and the guardrail down, by design, and both effects are real. Define the decision rule before the test runs: what magnitudes make this a ship and what makes it a stop. Deliverable: the rule, including the exchange rate you are using between the two quantities.
Approach
- Name the conflict precisely rather than calling it a balance: the reminder buys one week of an account returning and spends permission to contact that account, and permission is not renewable, so the two quantities are not comparable as percentage points.
- Put both sides into one unit before arguing about thresholds. Value the primary gain as incremental core-action weeks over the horizon; value an opt-out as the forgone email-driven active weeks over that account's remaining expected lifetime.
- Read opt-out as a stock, not a flow: accumulate it over the test, because a weekly opt-out rate that looks small is a cumulative curve that only ever rises within a fixed set of contactable accounts.
- Run long enough for the novelty to decay and use the late number in the trade: compare week-1 lift with week-4 lift, state the decay you observed, and refuse to price the decision on a week-1 read.
- Write the outcome as two numbers and a default action, including what happens when the result lands between them — hold and test a smaller volume increment rather than shipping on ambiguity.
Worked solution 30 min
- Measure the starting stock: current opt-out and unsubscribe rate per 1,000 contactable accounts, and the share of weekly active accounts whose session began from referrer_channel = 'email'.
- Estimate the two effects separately over four weeks: incremental weekly active accounts per 1,000 extra sends, and incremental cumulative opt-outs per 1,000 extra sends.
- Convert opt-outs into forgone active weeks: email-driven active weeks per account per year multiplied by remaining expected account lifetime, and set that against the incremental active weeks bought.
- Compare week-1 with week-4 lift, state the decay rate, and carry the week-4 figure into the trade.
- Write the rule as two thresholds plus a default action for the middle case.
Follow-up
- Opt-out is flat but week-over-week repeat rate falls. What is the most likely mechanism, and does it change the decision?
- You do not have a year of data to estimate remaining account lifetime. What do you substitute, and how do you keep the answer honest about that?
- Does randomising on account rather than on user change either the readout or the size of the test?
Gross revenue churn doubled with no cancellations behind it
Gross monthly revenue churn computed from fct_subscription_period doubled from 1.8% to 3.6% in one month. The support queue shows no rise in cancellations and renewals look normal. You have fct_subscription_period with subscription_id, account_id, period_start_utc, period_end_utc, mrr_cents, mrr_cents_constant_fx, seats_billed, period_status, change_reason and canceled_at_utc, plus dim_account. Decompose the 1.8-point rise into named mechanisms, size each in points of the headline, and state the remainder you cannot explain.
Approach
- Reconstruct the numerator row by row and group it by change_reason before arguing about causes. A mid-period plan or seat change closes the current period row and opens a new one, so any implementation that treats a closed row as lost revenue books upgrades, downgrades and seat changes as churn; the change_reason breakdown of the numerator makes that visible in one query.
- Check the recognition timestamp. Churn belongs at period_end_utc, because revenue continues until the period ends, not at canceled_at_utc when the button was pressed. Then count period_end_utc rows per month across thirteen months: annual cohorts concentrate their period ends in the month twelve months after they were signed, so a spike that repeats in the same month last year is seasonality in the book, not an event this month.
- Recompute the whole numerator and denominator on mrr_cents_constant_fx. If the constant-currency figure is materially flatter, the move is an exchange-rate translation and belongs nowhere in a churn narrative.
- Check period_status handling. Rows with period_status = 'past_due' are dunning, not cancellation; a status-based rule counts them as loss while a period-end rule does not, and a dunning backlog can move the number by itself.
- Express every mechanism in points of the headline, sum them, and print the residual explicitly next to the month-to-month standard deviation of the prior twelve months. A decomposition without a stated remainder is a story, not an accounting.
Follow-up
- Which of these mechanisms should be fixed in the metric definition and which should be reported as a genuine business fact?
- How would you present a month whose churn is dominated by an annual cohort anniversary without the audience concluding the business is deteriorating?
- What would you change so that an upgrade can never enter the churn numerator again, and how would you test that it worked?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
How do you handle multiple testing corrections when evaluating hundred…
How do you handle multiple testing corrections when evaluating hundreds of features simultaneously across different subsets?
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Explain a wide interval to a non-technical executive
A pricing change is under consideration. Your best estimate of its effect on trial-to-paid conversion is a 1.8pp drop, with a 95% interval from a 4.6pp drop to a 1.0pp rise, read from a geo holdout rather than a randomised test. An executive preparing a board slide asks you for 'the number'. You have ninety seconds and one slide, and the words confidence interval, p-value and significance are not usable with this audience. Deliver the slide headline, the single supporting line, and what you say aloud.
Approach
- Recognise what is being probed: whether you can carry uncertainty into a decision instead of either hiding it or hiding behind it. The generic answer promises to explain the interval in plain English; the strong one replaces the question 'what is the number' with 'across this range, where does the decision change'.
- Find the threshold before you draft anything. Ask what the pricing case assumes, then compute the conversion drop at which the higher price stops adding revenue: price uplift on the conversions kept against the revenue lost from conversions forgone. That single figure is what makes the range legible.
- Restate the estimate and both bounds in the unit the audience already reasons in. Convert percentage points into monthly first-paid conversions at current trial volume, then into mrr_cents_constant_fx, so the slide reads as money per month rather than as statistics.
- Place the range against the break-even and say which part of it sits on each side. If most of the range clears the threshold, that is a recommendation to proceed with a monitoring plan; if the range straddles it, that is a recommendation to narrow the range first.
- Name what would narrow it and what that costs in weeks, then give one recommendation with an explicit condition for revisiting it. Uncertainty stated without a next step is read as indecision and the midpoint gets used anyway.
Follow-up
- The executive says to give the midpoint and they will manage the risk. What do you do?
- How does the slide change if the interval were a 4.6pp to 0.2pp drop, with no positive outcomes in range?
- Why is a geo holdout the credible read here rather than the attributed channel numbers you already have?
Handle a request for numbers supporting a decision already made
A senior leader has already decided to sunset a plan tier and asks you for the analysis showing it is the right call. Accounts on that tier carry 6% of MRR at constant FX and have the highest licensed-seat utilisation in the book. The leader's support matters to your next review cycle, and the decision is being presented in four days. Deliver what you produce, what you decline to produce, and the exact sentence you will say in the meeting where the number appears on a slide.
Approach
- Recognise what is being probed: whether you can find the legitimate request inside an illegitimate framing instead of either complying or refusing on principle. The generic answer promises to push back; the strong one produces something genuinely useful and states its limits in the room, without ambushing anybody.
- Separate the decision from the justification. Sunsetting the tier may be correct for reasons the data does not hold, such as support cost, roadmap surface area or sales motion. What you decline is a one-sided document. What you produce is the case read both ways, which also happens to be more useful to the leader.
- Build the symmetric analysis: MRR at risk at constant FX, the share of affected accounts with a plausible migration path given seats_licensed and billing_term, the recovery rate assumed for that migration and where it came from, and the downside case in which high-utilisation accounts treat the sunset as a reason to re-evaluate the vendor entirely.
- Surface the inconvenient fact privately and early. The highest seat utilisation in the book is a retention signal, and the leader should hold it before the room does, so they can incorporate it rather than be caught by it.
- Agree the meeting sentence in advance with the leader, so that nobody is surprised. Something to the effect that the tier is 6% of MRR and its accounts are the most heavily used in the book, and that the case for sunsetting rests on cost and focus rather than on revenue. That is true, it supports the decision on its real grounds, and it stops the deck claiming the numbers endorse it.
- Decide your own line before you need it: what you will not put your name to, and that the route if asked anyway is your own manager rather than a confrontation in the meeting.
Follow-up
- The deck circulates with your analysis included and the downside case removed. What do you do, and by when?
- What changes if the honest analysis says the sunset is clearly the wrong call?
- How do you write the same memo when the leader is your skip-level and the meeting is tomorrow?
- 01
How do you handle multiple testing corrections when evaluating hundreds of features simultaneously across different subsets?
- 02
A pricing change is under consideration. Your best estimate of its effect on trial-to-paid conversion is a 1.8pp drop, with a 95% interval from a 4.6pp drop to a 1.0pp rise, read from a geo holdout rather than a randomised test. An executive preparing a board slide asks you for 'the number'. You have ninety seconds and one slide, and the words confidence interval, p-value and significance are not usable with this audience. Deliver the slide headline, the single supporting line, and what you say aloud.
- 03
A senior leader has already decided to sunset a plan tier and asks you for the analysis showing it is the right call. Accounts on that tier carry 6% of MRR at constant FX and have the highest licensed-seat utilisation in the book. The leader's support matters to your next review cycle, and the decision is being presented in four days. Deliver what you produce, what you decline to produce, and the exact sentence you will say in the meeting where the number appears on a slide.
Is this an official Voleon interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Voleon. Rounds and questions reflect what candidates have reported, not a process Voleon has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews at The Voleon Group?
The technical interviews are rigorous, focusing heavily on foundational statistics, probability, and practical coding. While the questions are grounded in real-world data tasks rather than abstract algorithmic puzzles, interviewers expect complete, production-ready solutions and deep conceptual clarity.
PracHub interview research ↗How much preparation time should I plan for?
Most successful candidates dedicate between four to six weeks of focused preparation. This time should be split evenly between refreshing core statistical theory, practicing live SQL and Pandas data manipulation, and working through probabilistic reasoning problems.
PracHub interview research ↗What distinguishes successful candidates from those who are rejected?
Successful candidates excel at communicating their thought process aloud, structuring ambiguous problems methodically, and validating their assumptions before writing code. Conversely, candidates who rush to code without verifying edge cases or who struggle to explain the underlying math of their models rarely advance.
PracHub interview research ↗What is the typical interview timeline from initial screen to offer?
The process typically spans several weeks to a couple of months, beginning with an HR screen, moving through technical phone screens or Hackerrank assessments, and culminating in virtual or on-site multi-round loops followed by reference checks.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22