At Product Madness, a Data Scientist is at the absolute center of player experience, game economy, and growth strategy. Operating within the Product & Design and Growth divisions, you will transform massive streams of player telemetry data into actionable insights that directly shape how millions of users interact with top-tier social casino and mobile games. This is not a purely advisory role; your models, statistical analyses, and data pipelines will directly power live-ops, personalization engines, and user acquisition strategies.
The impact of this position cannot be overstated. With a massive global player base, even minor optimizations in player retention, in-game economy balancing, or ad targeting can lead to significant shifts in business performance. You will collaborate closely with product managers, game designers, and software engineers to design sophisticated A/B tests, build predictive machine learning models, and uncover deep behavioral patterns.
To succeed in this role, you must possess a unique blend of core programming discipline, rigorous statistical foundations, and sharp product intuition. values candidates who do not just run algorithms blindly, but who can think critically about the underlying player behavior and business dynamics that the data represents.
HR Screening Call
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Test Task
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Python Coding Interview
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Machine Learning Session
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Statistics Interview
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
PracHub editorial advice for the preparation topics above.
Booking revenue at gross, on purchase day, for cohorts of different ages.
Revenue matures: the platform fee is known immediately, but refunds and chargebacks land over weeks, and later purchases keep accruing to the same cohort. Comparing a two-week-old cohort with a one-year-old cohort therefore compares two different maturities, and any trend chart built that way slopes in the direction of cohort age. Quote cohort revenue at a fixed age, and exclude cohorts that have not reached it rather than extrapolating them.
Randomising players individually for a change that acts on a shared pool.
Matchmaking parameters, queue populations, tradeable-item supply and event leaderboards are shared resources: a treated player is matched against a control player, and treated supply lands in the control arm's market. The treatment leaks across arms, so the measured difference understates or reverses the true effect, and the control arm is no longer a clean counterfactual. Cluster by region and mode, or switchback on time slices with a burn-in long enough to clear carryover in the shared state.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
SQL that silently fans out on a one-to-many join
State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you handle highly imbalanced datasets when training a classific…
How do you handle highly imbalanced datasets when training a classification model for VIP player identification?
Approach
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- What would you monitor after launch to know the model is still valid?
- Where could label leakage enter this setup?
How would you evaluate the performance of a recommendation engine desi…
How would you evaluate the performance of a recommendation engine designed to suggest in-game store packages to players?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Rebuild running balances and find the first ledger divergence
Using the same ledger frame plus balance_after (int64), rebuild each player's running balance per currency_code by accumulating delta_amount in (occurred_at, ledger_id) order, then report the first row per (player_id, currency_code) where your computed balance disagrees with balance_after. You may not use cumsum, groupby().cumsum() or expanding(); accumulate explicitly. Assume each group starts from a balance of zero. Reversals and corrections do count toward the balance. Return player_id, currency_code, ledger_id, computed_balance, balance_after and diff, sorted by absolute diff descending.
Approach
- Sort once on (player_id, currency_code, occurred_at, ledger_id). The ledger_id tiebreak is load-bearing: entries routinely share occurred_at to the second, and an unstable order manufactures divergences that move between runs.
- Accumulate in one pass over itertuples with a dict keyed by (player_id, currency_code), resetting to zero when the key changes. One pass is O(n) and beats any reshape; the ban on cumsum is about showing the accumulation, not about speed.
- Record the first index per key where computed != balance_after and then stop tracking that key. After a break every later row inherits the same offset, so reporting all of them buries the single row that matters.
- Keep reversals and corrections in the accumulation. They are genuine balance movements, which is the exact opposite of the rule for flow aggregates, and mixing the two rules up is the common error in this domain.
- Sort the output by abs(diff) descending. A one-unit diff is usually a rounding artefact somewhere upstream; a six-figure diff is a lost grant, and the ordering puts the second kind on the first screen.
Worked solution 25 min
- Sort on the four-column key and reset the index so positional recording is unambiguous.
- Walk the frame once, carrying a dict of running balances and a set of keys already flagged, appending a record the first time a key disagrees.
- Assemble the flagged records into a frame, compute diff = computed_balance - balance_after, and sort by its absolute value descending.
- Spot-check by re-accumulating one flagged group by hand from its first row and confirming the break lands on the reported ledger_id.
Follow-up
- Your check assumes each group starts at zero. How would you distinguish a player whose early history was trimmed by a retention policy from one who genuinely started at zero?
- Divergence appears on exactly one currency and only after a specific date. What is your next query?
- How would you run this as a daily monitor over hundreds of millions of rows without rebuilding every balance from the beginning of time?
Explain the difference between mutable and immutable objects in Python…
Explain the difference between mutable and immutable objects in Python and how this impacts memory management.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Write a Python function to find the first non-repeating character in a…
Write a Python function to find the first non-repeating character in a string and analyze its time and space complexity.
Approach
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Measure install-to-first-match conversion within the first twenty-four hours
dim_player has player_id, install_ts, install_platform, is_test_account. fct_match_participant has player_id, started_at, result. Report conversion by UTC install date and install_platform: cohort size, and the share of that cohort with at least one match where started_at < install_ts + 24 hours and result <> 'abandon'. Exclude test accounts. Maturity is a property of the whole cohort, not of individual players: an install date is reportable only once every install on it has had its full 24 hours, and the last install on a date lands at 23:59:59 UTC. Report the 30 most recent install dates that satisfy that. Every install in a reported cohort stays in the denominator whether or not it converted.
Approach
- Build the cohort from dim_player alone so the denominator is fixed before any match data is consulted: is_test_account = false and date(install_ts) BETWEEN current_date - 31 AND current_date - 2. The upper bound is the maturity rule. A date D is complete only from D + 2 days onward, because an install at D 23:59:59 UTC does not reach 24 hours of age until D+1 23:59:59, so current_date - 2 is the newest date that is safe at any hour of the run.
- Do not express maturity as a row-level filter. install_ts < now() - interval '24 hours' looks equivalent and is not: on a run at 14:00 it keeps the players who installed before 14:00 yesterday and drops the rest, so yesterday's cohort is a partial slice whose size grows through the day and whose composition is skewed toward the hours that happen to precede your run. Evening installers are not early-morning installers, so the rate moves with clock time rather than with the product. Filter whole dates.
- Left join fct_match_participant with the started_at and result conditions in the ON clause, not the WHERE clause. In the WHERE clause those conditions convert the left join into an inner join and the denominator quietly becomes players who already played.
- Count distinct converting player_id rather than matching rows, since a player who plays four matches in the first hour must count once.
- Group by date(install_ts) and install_platform, and report cohort_size beside the rate so a 100 percent rate on a cohort of three is visible as such.
- State the identity key in the output. This counts player_id, so a reinstall that mints a new account appears as a fresh, unconverted install and depresses the rate on platforms where reinstalls are common.
Worked solution 15 min
- Write the cohort CTE from dim_player with the test-account filter and date(install_ts) BETWEEN current_date - 31 AND current_date - 2, then count rows per (install_date, install_platform) to fix the denominator.
- Add the left join with all match conditions in ON, selecting only the match-side player_id.
- Aggregate: COUNT(*) as cohort_size and COUNT(DISTINCT m.player_id) as converted_players.
- Compute converted_players::numeric / cohort_size and confirm no row has converted_players greater than cohort_size.
- Size the truncation you avoided: count installs dated current_date - 1 that satisfy install_ts < now() - interval '24 hours' against the full count for that date. The shortfall is what a row-level maturity filter would have reported as a complete cohort.
Follow-up
- result <> 'abandon' silently drops any row where result is NULL. Is that the behaviour you want, and how would you write the predicate so the intent is explicit?
- Conversion is flat but daily cohort size doubled after a paid campaign launched. What cut do you run before concluding the campaign brought good installs?
- How does the number change if you anchor the 24 hours to first session start instead of install_ts, and which anchor would you defend?
How would you optimize a memory-heavy data processing script in Python…
How would you optimize a memory-heavy data processing script in Python when handling large player log files?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What is statistical power, and how do you calculate the required sampl…
What is statistical power, and how do you calculate the required sample size for an experiment with low baseline conversion rates?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
How would you design an A/B test for a new in-game monetization featur…
How would you design an A/B test for a new in-game monetization feature when player activity is highly skewed?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
Describe how you would handle multiple testing problems (e.g., family-…
Describe how you would handle multiple testing problems (e.g., family-wise error rate) when evaluating several game variants simultaneously.
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you do if you could not randomise at all?
- How would you handle interference between treated and control units?
Recover a causal effect from a catch-up grant threshold
Halfway through a seasonal event, every player below pass tier 20 was automatically granted 500 premium currency; players at tier 20 or above received nothing. The grant was not randomised, and tier is visible to players throughout. Using fct_currency_ledger (source_system, delta_amount, occurred_at), the midpoint tier snapshot, and fct_iap_transaction, estimate the effect of the grant on net revenue over the remaining three weeks. State the design, the bandwidth and functional form, the estimand you can defend, and the single test that can invalidate the whole thing.
Approach
- Identify the design: assignment is a deterministic step function of a known running variable, midpoint tier, at a known cutoff of 20, so this is a sharp regression discontinuity. The comparison is between players just below and just above 20, not between the granted and ungranted populations as a whole.
- Specify estimation: local linear regression on each side of the cutoff with a triangular kernel and an MSE-optimal bandwidth, reporting bias-corrected robust confidence intervals rather than conventional ones, which undercover at the MSE-optimal bandwidth. Report the estimate at half and double the bandwidth as sensitivity. Avoid high-order global polynomials, which are known to manufacture discontinuities at the boundary.
- Handle the discreteness. Tier is an integer with few mass points near the cutoff, so treat the running variable as discrete, cluster standard errors by tier value, and be explicit that the effective sample is the handful of tiers inside the bandwidth, not the row count.
- State the estimand honestly: a local average treatment effect at tier 20 for players sitting at the cutoff at the midpoint. It supports the decision 'should the cutoff move' and does not support 'what would this grant do for a tier-5 player'.
- Run the invalidating test: manipulation of the running variable. Tier is visible and the grant was announced, so players had both the motive and the means to stall just below 20. Test the density of the running variable at the cutoff for bunching, and test that pre-determined covariates are continuous there: install cohort, platform, pre-midpoint spend, matches played before the announcement. Bunching or a covariate jump kills the design, and a donut specification excluding the tiers adjacent to the cutoff is a diagnostic, not a repair.
- Check for co-located confounds and mechanism. Anything else that switches at tier 20, a reward unlock or a shop gate, is inside the estimate, because RD identifies the combined effect of everything that changes at the cutoff. Then use fct_currency_ledger to confirm the granted 500 was actually spent within three weeks; a grant that sits unspent cannot have moved revenue through the claimed mechanism.
Worked solution 40 min
- Build the analysis table: midpoint tier as the running variable, a treated indicator for tier below 20, and net revenue over the following three weeks from fct_iap_transaction at a fixed maturity.
- Fit local linear regressions either side of 20 with a triangular kernel at the MSE-optimal bandwidth, and report the bias-corrected robust interval.
- Repeat at half and double the bandwidth, and at placebo cutoffs of 15 and 25.
- Run the density test at the cutoff and the continuity tests on install cohort, platform, pre-midpoint spend and pre-announcement matches.
- Measure the sink volume of the granted currency in fct_currency_ledger within three weeks, split either side of the cutoff, as the mechanism check.
- Write the estimate as a LATE at tier 20, with the validity tests attached to it rather than in an appendix.
Follow-up
- The density test shows a spike just below tier 20. What can you still estimate, and what would you refuse to report?
- Would you randomise the grant next season, and what would you have to give up to run it?
- The estimate at the cutoff is +$0.42 per player. How would you turn that into a recommendation about where the cutoff should sit?
Sessions per player dropped the week a client version shipped
Sessions per active player fell 31% the week a client version rolled out, while matches per active player and total play seconds are flat. fct_session has session_id, player_id, session_start_ts, session_end_ts, duration_seconds, app_version, matches_started, ended_reason. A growth lead has filed this as an engagement regression. Determine whether player behaviour changed or the session boundary did, using only these columns, and state the evidence that would settle it either way.
Approach
- Start from the conservation check that is already in the prompt: if matches per active player and total play seconds are both flat while session count fell, the same activity is being partitioned into fewer containers. Genuine disengagement would move at least one of those two.
- Verify the implied arithmetic. If sessions per player fell 31%, matches per session must have risen by 1/0.69 minus 1, about 44.9%, for matches per player to stay flat. Compute it. A match that lands close to that figure is strong evidence of re-partitioning rather than behaviour.
- Split the same calendar day by app_version. Holding the date fixed removes weekday, release-calendar and server-side effects, so a step difference between versions on one day localises the change to the client. Note the caveat aloud: version adoption is not random, since fast updaters differ from slow ones, so this is corroboration, not proof.
- Get mechanism evidence from the gap distribution. For each player, compute session_start_ts minus the lagged session_end_ts, and histogram those gaps by app_version. If the client's idle timeout lengthened, the old version has mass in the band between the old and new timeouts and the new version has none, because those gaps now sit inside a single session. The cut point names the new timeout.
- Handle the null trap explicitly: duration_seconds is null whenever session_end_ts is null, which is what a crashed client leaves behind. An average over duration_seconds silently drops those rows, so report the null rate by app_version and by ended_reason alongside any duration statistic.
- Deliver the verdict with a restatement plan: which historical series must be recomputed on a consistent boundary, and which metrics, such as matches per active player, were never affected and can be compared across the change.
Follow-up
- Which downstream metrics in this product are defined on sessions rather than on players or matches, and what does each of them now mis-state?
- How would you backfill a comparable series across the boundary change, and what would you refuse to backfill?
- If ended_reason='unknown' also rose on the new version, does that strengthen or weaken your conclusion?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Have two ready. In one, the data was on your side and you had to move someone who outranked you. In the other, the pushback was correct and you changed position. The second is the harder story and it lands better, because it shows you separate being right from being attached to an answer. Name the person's actual objection.
Report a contaminated switchback when the ship call is Friday
A switchback test of a container drop-rate change ran Monday to Sunday on 12-hour slices, randomising the whole region-mode population per slice. On Wednesday a limited-time event went live to everyone, changing both the currency faucet and store traffic for the remaining slices. Treatment and control slices are no longer exchangeable. The change is scheduled to ship Friday and the director wants a number. Give what you say, what you can still salvage from the data, and what you ask for.
Approach
- Say the full-week estimate is unusable before saying anything else, and give the reason in one sentence a non-analyst can repeat: the world changed halfway through, so the two halves of the test measure different games and pooling them averages across that difference.
- Test salvageability rather than assuming it away. If slice assignment is balanced within the pre-event window and that window holds enough slices for a usable standard error, you have a shorter and noisier but uncontaminated estimate. Report it with the widened interval and an explicit statement of the power you lost. Do not pool across the event boundary, and do not run the post-event slices as a second test with the event as a covariate, because the event was not randomised.
- Check carryover independently of contamination, because they fail for different reasons. A drop-rate change leaves players holding containers and currency, so treatment effects persist into the following control slice unless the burn-in exceeds the time it takes a typical player to open what they received. State the burn-in you used and what evidence supports it.
- Give the decision a path rather than a blocker. Offer a cluster-randomised staged rollout by region and mode with the guardrail set pre-registered before launch, which gets the change shipped on Friday while still producing a clean estimate.
- Fix the cause, not the instance. The live-ops calendar was a discovery here and should be a design input, checked against the planned slice schedule before a test is launched.
- Decide in advance what you will do if the director insists on a single number anyway, and prefer publishing the clean partial estimate with its interval over publishing the contaminated one with a caveat.
Follow-up
- The pre-event window has six slices. What standard error does that give you, and is it worth reporting at all?
- How would you choose the burn-in length for the rerun, using data rather than judgement?
- The director ships it anyway on Friday. What do you measure afterwards, and what can that measurement honestly support?
Allocate one analyst-week across three competing escalated requests
Three requests land in the same week and you have five working days. The economy team wants a sink-to-faucet audit per currency after a faucet change shipped ten days ago. Live-ops wants a readout on an event that ends Friday, because the next event is configured from it. Acquisition wants 90-day net revenue per install by channel for a budget meeting in three weeks, and two of the channels launched six weeks ago. All three owners have escalated. Give your allocation, the reasoning you give each owner, and what you refuse or defer.
Approach
- Sort by decision deadline and by reversibility rather than by escalation volume. The event readout is perishable because the population and the live-ops configuration that produced it stop existing on Friday and the next event's config depends on it. The budget meeting is three weeks out. The economy audit has no external deadline but a compounding cost.
- Kill the part that cannot be done correctly at any effort level, and kill it in a ten-minute conversation rather than four days of work. Net revenue per install at 90 days requires cohorts that have reached 90 days of maturity; channels that launched six weeks ago have not, and extrapolating them produces a number that will slope with cohort age. The honest deliverable is matured channels only, with the immature ones listed as excluded and dated for when they qualify.
- Split the economy request into the decision-relevant core and the rest. One day gets the sink-to-faucet ratio per currency_code for the weeks before and after the faucet change, with reversals, transfers and cs_grant excluded, plus the balance percentile curve. A ratio below 1 sustained means balances are accumulating and premium shortcuts will stop selling, which is worth knowing this week. The full per-source audit can wait.
- Give the event readout the largest block, because it is the one with a hard expiry and a downstream configuration decision. Scope it to a decision memo, not a dashboard.
- Publish the allocation in one place with a one-line reason per item, so any escalation argues with the reasoning rather than with you, and the owners can see each other's deadlines.
- Hold back roughly one day. Something breaks most weeks, and an allocation with no slack fails in a way that damages all three commitments instead of one.
Follow-up
- The acquisition owner says a rough number is better than nothing for a budget meeting. What exactly do you give them?
- How would you decide whether the economy audit is genuinely urgent rather than merely important?
- Two weeks of this pattern in a row. What structural change do you propose, and to whom?
Turn 'engagement is down' into a scoped answerable brief
A studio lead messages: engagement is down, can you look into it. You have dim_player, fct_session, fct_match_participant and the release calendar. Daily core-loop players, defined as distinct players per UTC day with at least one match reaching a terminal result other than abandon, is down 6% week over week, and a content release landed nine days ago. You have 30 minutes before a standup. Produce the scoping questions you would ask, the first three cuts you would run, and the one-paragraph brief you send back before doing deeper work.
Approach
- Pin the metric and the comparison before touching data. Ask which number the lead actually saw and over what window, because a week-over-week read nine days after a release is measuring post-release decay by construction, and that alone may be the whole answer.
- Ask the decision question, not more metric questions: what would the lead do differently if this turns out to be new players versus returning, one platform versus all, one region versus global. Scope follows the decision, and an investigation with no decision attached should be declined or deferred.
- Run three cheap cuts that split the space rather than confirm a hunch. First, new versus existing by install cohort age, which separates an activation problem from a retention problem. Second, platform crossed with app_version, where a bad build shows as a concentration of fct_session.ended_reason = 'crash' and truncated duration_seconds. Third, server_region, where an infrastructure incident shows as elevated match abandon rate and p95 matchmaking wait rather than as fewer app opens.
- Re-baseline against the matched day in the previous release cycle instead of against last week, so the comparison is not dominated by the release calendar.
- Send a brief that states what you confirmed, what you ruled out, the current best explanation with its confidence, and the size of the next block of work with the question it would close.
Follow-up
- The crash concentration is on one device_model at one app_version. What do you send, to whom, and how urgently?
- How would you tell a genuine drop apart from an instrumentation change that altered which sessions get logged?
- If all three cuts come back flat, what is your fourth cut and why that one?
- 01
A switchback test of a container drop-rate change ran Monday to Sunday on 12-hour slices, randomising the whole region-mode population per slice. On Wednesday a limited-time event went live to everyone, changing both the currency faucet and store traffic for the remaining slices. Treatment and control slices are no longer exchangeable. The change is scheduled to ship Friday and the director wants a number. Give what you say, what you can still salvage from the data, and what you ask for.
- 02
Three requests land in the same week and you have five working days. The economy team wants a sink-to-faucet audit per currency after a faucet change shipped ten days ago. Live-ops wants a readout on an event that ends Friday, because the next event is configured from it. Acquisition wants 90-day net revenue per install by channel for a budget meeting in three weeks, and two of the channels launched six weeks ago. All three owners have escalated. Give your allocation, the reasoning you give each owner, and what you refuse or defer.
- 03
A studio lead messages: engagement is down, can you look into it. You have dim_player, fct_session, fct_match_participant and the release calendar. Daily core-loop players, defined as distinct players per UTC day with at least one match reaching a terminal result other than abandon, is down 6% week over week, and a content release landed nine days ago. You have 30 minutes before a standup. Produce the scoping questions you would ask, the first three cuts you would run, and the one-paragraph brief you send back before doing deeper work.
Is this an official Product & Design interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Product & Design. Rounds and questions reflect what candidates have reported, not a process Product & Design has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the Python coding interview?
It is highly technical but focused on fundamentals. You will not necessarily face complex dynamic programming questions typical of FAANG software engineering rounds, but you must demonstrate clean code, efficient data structure usage, and strong algorithmic logic without relying on external libraries.
PracHub interview research ↗What is the company culture like within the data team?
Candidates consistently report that the data science team is highly knowledgeable, collaborative, and friendly. The company maintains a strong spirit of innovation, though the interview process itself is highly structured and demanding.
PracHub interview research ↗How should I prepare for the final round with the Director or Hiring Manager?
Do not treat this as a purely casual chat, even if it is pitched as one. Be prepared for deep analytical questions, abstract machine learning design scenarios, and logical problem-solving exercises. They want to see how you think on your feet when data is limited.
PracHub interview research ↗What is the typical timeline for the interview process?
The timeline can vary. While HR is typically prompt in initial stages, the overall process can take several weeks due to the high number of technical rounds and scheduling across multiple senior stakeholders.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22