Epic Games · Data Scientist
Updated · 2026-09-22

Epic Games Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

The role of a Data Scientist at Epic Games is crucial in shaping the gaming experience and driving business strategies. As a Data Scientist, you will analyze vast amounts of data generated by players across Epic's games and platforms, such as Fortnite and the Epic Games Store. This position not only influences product development but also enhances user engagement and retention by delivering actionable insights into player behavior, preferences, and trends.

Scope your preparation by the data you would actually touch, because the title will not tell you. A seat that lives in event logs and weekly readouts rewards fluency in aggregation and metric definitions; a seat that owns a model in production rewards fluency in train/serve skew, retraining cadence and drift monitoring. The fastest way to find out which one you are interviewing for is to ask what the team shipped last quarter and what it gets paged about.

Epic Games candidates report 2 rounds · ≈ 2-4 weeks. The stages below are what candidates describe, not a published process.

Price cosmetics without cannibalising existing sinksCompute expected cost of pity-timer tablesDesign matchmaking tests that survive interference

33 min read

Practice 15 Data Scientist prompts
15Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

The role of a Data Scientist at Epic Games is crucial in shaping the gaming experience and driving business strategies. As a Data Scientist, you will analyze vast amounts of data generated by players across Epic's games and platforms, such as Fortnite and the Epic Games Store. This position not only influences product development but also enhances user engagement and retention by delivering actionable insights into player behavior, preferences, and trends.

Your work will directly impact how teams across the company make decisions, ranging from game design to marketing strategies. By applying statistical analysis, machine learning, and data visualization techniques, you will uncover patterns that can lead to innovative game features and improved player satisfaction. This role is both challenging and rewarding, offering the opportunity to work on complex datasets while collaborating with cross-functional teams in a dynamic gaming environment.

01

Take-Home Challenge

reported

The clock is part of the test. Three to six hours is not enough to do everything the dataset supports, so the submission mostly reveals how you spend a fixed budget against an open question. A reviewer sees which paths you took and, by absence, which you abandoned. Work that runs out of time inside the analysis ships a thin conclusion, while work that cuts scope early protects the last hour for writing. The most reliable way to lose here is to leave the scoping decision implicit, so it reads as something you missed rather than something you chose.

What to demonstrate

  • Whether the scope you settled on is presented as a decision with a reason, rather than left for the reader to infer from what is missing
  • Whether the depth of the work is consistent with the stated time budget, instead of several half-finished directions left open
  • Whether the closing section reads as something written on purpose rather than assembled from whichever cells survived

How to prepare

  • Run a timed rehearsal on a public dataset with a hard stop, holding the final sixty minutes for writing no matter where the analysis has got to
  • Before opening the data, list the questions it could plausibly answer, pick one, and keep the discarded ones as a short note on what you did not attempt and why
  • Commit a one-line finding after each analysis step so the writeup is assembled from recorded results rather than from memory at midnight
PracHub interview research ↗
02

Technical Interviews

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Comparing retention or spend across progression stages reached.

Reaching level 20 requires having survived long enough to reach level 20, so grouping by max level attained conditions on the outcome you are trying to explain. Every correlate of playing longer then looks causal, and 'players who join a guild retain better' is the canonical version of this error. The defensible framing fixes a cohort at a common age or a common exposure point and measures forward from there, or models the hazard with progression as a time-varying covariate.

02

Treating the account as the person, and the calendar week as neutral time.

Reinstalls, multi-platform play and shared accounts split or merge humans relative to player_id, which inflates install counts and deflates retention denominators in ways that correlate with platform and campaign. Meanwhile daily actives spike on release day and decay over the following week, so a week-over-week read aliases the content calendar. Anchor comparisons to time since release or to matched days in the release cycle, and report identity-sensitive metrics with an explicit statement of which identity key was used.

03

Crediting a treatment for regression to the mean

Selecting a group because it is extreme (lowest-engagement users, accounts having their worst month, the bottom decile of a score) moves that group's expected next-period value back toward the average even under no treatment, by exactly as much as the selecting measure is imperfectly correlated with its own later value. Compare against units that met the same selection rule and went untreated, or use two pre-periods so the bounce-back is visible before the intervention starts. A pre-post number on a group chosen for being extreme measures the selection rule, not the treatment.

04

Dropping rows with missing values without naming the mechanism

Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

12 technical prompts3 include a worked solution

Explain your thought process while implementing a specific algorithm.

medium
machine learning and modelling

Explain your thought process while implementing a specific algorithm.

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • Where could label leakage enter this setup?

What metrics would you use to evaluate the performance of a model?

medium
machine learning and modelling

What metrics would you use to evaluate the performance of a model?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Pick an evaluation metric that matches the cost of each error type, not a default.
  3. Say how the offline result would be validated online before it is trusted.
Follow-up
  • Where could label leakage enter this setup?
  • What would you monitor after launch to know the model is still valid?

Discrete hazard of churn indexed by matches played

hardWorked solution
survival analysiscensoringchurn

Given matches (player_id, started_at) for one install cohort, sessions (player_id, session_start_ts), and a data cutoff timestamp, define churn at match k as no session starting in (t_k, t_k + 14 days], where t_k is the start of that player's k-th match. Build the discrete hazard h(k) and the survival curve S(k) for k = 1 to 50. A player-index enters the risk set at k only if the player has not already churned before k and t_k is at or before cutoff minus 14 days. Return k, risk_set, churned, h and S, and name the two largest hazard spikes.

Approach
  1. Build the match index yourself: sort matches by (player_id, started_at) and set k = groupby('player_id').cumcount() + 1. Never take a stored index without verifying it is dense per player, because one gap shifts every downstream k for that player.
  2. Decide churn per (player, k) without a cross join. Sort each player's session timestamps once and use np.searchsorted to ask whether any session falls in the half-open window (t_k, t_k + 14d]. A merge of matches to sessions is quadratic in the busiest players and is where this exercise usually dies.
  3. Keep observability separate from the event. A player-index is usable at k only when t_k is at or before cutoff minus 14 days; a player whose 30th match was yesterday contributes to k = 1 through 29 and then leaves. This is right-censoring on the match axis, and dropping those players outright instead of truncating them biases the entire curve toward whoever plays fastest.
  4. Compute h(k) = churned_at_k / risk_set_k and S(k) as the cumulative product of (1 - h(j)) for j up to k. The product form is what makes the censoring handling count; a naive share of the original cohort still alive divides by a denominator containing players who could never have been observed at k.
  5. Read the spikes against the loop rather than the calendar: a bump at a given k usually maps to a difficulty wall, the end of a tutorial reward track, or the point where a faucet stops paying. Name the k values, then name the query that would confirm each one, rather than asserting a cause.
Worked solution 40 min
  1. Derive t_k per player by sorting and using cumcount, then truncate each player's indices at the last k satisfying t_k at or before cutoff minus 14 days.
  2. For each surviving (player, k), use a per-player searchsorted over session timestamps to decide whether a session falls in the 14-day window, and take each player's first churn index as their exit.
  3. Aggregate risk_set and churned counts by k for k = 1 to 50, then compute h(k) and the cumulative-product S(k).
  4. Rank h(k) and report the two largest, with the k values and the local shape around them.
EXPECTED RESULTA 50-row frame of k, risk_set, churned, h and S, with risk_set non-increasing in k, S non-increasing in k, and every churned player counted at exactly one k.
Follow-up
  • Someone reports that players who reach match 50 retain four times better. Rewrite that as a claim you would be willing to defend, and say what it costs you to do so.
  • Your definition retires a player who keeps logging in but stops playing matches. Is that churn? What does including or excluding them do to the curve?
  • Add a time-varying covariate for whether the player was in a party at match k. What changes about the estimator, and what can you now claim that you could not before?

For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Design one test end to end on paper
  • Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
  • Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
  • State in advance what you will do if the primary metric is flat while a secondary metric is significant.

Deliverable: A one-page test design with a decision rule written before launch.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Power arithmetic until it is automatic
  • Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
  • Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
  • Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.

Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.

Practice prompt ↗Practice prompt ↗
03Variance and the unit-of-analysis problem
  • Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
  • Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
  • Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.

Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.

Practice prompt ↗Practice prompt ↗
04Validity threats you can actually test for
  • Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
  • Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
  • Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.

Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05When randomization is not available
  • Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
  • Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
  • List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.

Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.

Practice prompt ↗Practice prompt ↗
06The readout query
  • Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
  • Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
  • Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.

Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.

Practice prompt ↗Practice prompt ↗
07Present it to someone who will not read the appendix
  • Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
  • Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
  • Rewrite your opening line so the recommendation lands before any methodology.

Deliverable: A one-page readout whose first line is the recommendation.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.

Describe a time when you had to collaborate with a difficult team memb…

medium
behavioural and stakeholder questions

Describe a time when you had to collaborate with a difficult team member.

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Pick a story where you drove the decision, not one where you observed it.
  3. Close with what you would do differently, concretely.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Defend an unpopular readout on a live-ops event

medium
stakeholder pushbackrevenue maturitypull-forwardeconomy health

A two-week live-ops event closed with gross bookings up 18% against the prior week. Your readout says it added no net revenue: summed price_usd_net in fct_iap_transaction over the 28 days spanning the event is flat against a matched window one release cycle earlier, and fct_currency_ledger shows premium-currency sinks falling while faucets from source_system = 'pass_tier' rose. The event owner disputes the readout in a review with their director present. Give the argument you would make in that room, and name the one piece of evidence you would concede.

Approach
  1. Separate the disputed claim from yours. Gross bookings up 18% is true and you are not contesting it. Restate your finding precisely: net of platform fee and of refunds recognised to date, over a window long enough to contain the pull-forward, the number does not move.
  2. Present the decomposition rather than the aggregate, because the aggregate is what the room is already arguing about. Split the delta into active player-days, paying conversion, and net spend per converting player-day. If conversion is flat and spend per payer spiked inside the event window then fell below baseline after, that is re-timing, not new demand.
  3. Evidence the re-timing at player grain: take the player_ids who purchased during the event and chart their purchase timing in the four weeks before and after. A pulled-forward purchase shows as a deficit on the same accounts, not as a smaller cohort.
  4. Bring the ledger in as the second, independent signal. Premium sinks fell while pass_tier faucets rose, so granted currency was banked rather than spent. Report the sink-to-faucet ratio per currency and the balance percentile curve, and exclude is_reversal rows and cs_grant so a correction cannot masquerade as a faucet.
  5. State the limit of the design out loud before the owner does. There was no holdout, and a matched window still aliases the release calendar. Say which part of your conclusion is robust to that and which is not.
  6. Close on a decision and a fix: run the next instance with a region-clustered holdout or a staggered start, so the argument is settled by design rather than by seniority.
Follow-up
  • The event owner says the banked currency will be spent next month, so the sink drop is timing too. How would you test that claim rather than argue about it?
  • If you could add exactly one instrumentation change before the next event, what would it be and what question does it close?
  • What would you have said if the net number had been up 3% with an interval spanning zero?

Allocate one analyst-week across three competing escalated requests

medium
prioritisationcohort maturitystakeholder managementdecision deadlines

Three requests land in the same week and you have five working days. The economy team wants a sink-to-faucet audit per currency after a faucet change shipped ten days ago. Live-ops wants a readout on an event that ends Friday, because the next event is configured from it. Acquisition wants 90-day net revenue per install by channel for a budget meeting in three weeks, and two of the channels launched six weeks ago. All three owners have escalated. Give your allocation, the reasoning you give each owner, and what you refuse or defer.

Approach
  1. Sort by decision deadline and by reversibility rather than by escalation volume. The event readout is perishable because the population and the live-ops configuration that produced it stop existing on Friday and the next event's config depends on it. The budget meeting is three weeks out. The economy audit has no external deadline but a compounding cost.
  2. Kill the part that cannot be done correctly at any effort level, and kill it in a ten-minute conversation rather than four days of work. Net revenue per install at 90 days requires cohorts that have reached 90 days of maturity; channels that launched six weeks ago have not, and extrapolating them produces a number that will slope with cohort age. The honest deliverable is matured channels only, with the immature ones listed as excluded and dated for when they qualify.
  3. Split the economy request into the decision-relevant core and the rest. One day gets the sink-to-faucet ratio per currency_code for the weeks before and after the faucet change, with reversals, transfers and cs_grant excluded, plus the balance percentile curve. A ratio below 1 sustained means balances are accumulating and premium shortcuts will stop selling, which is worth knowing this week. The full per-source audit can wait.
  4. Give the event readout the largest block, because it is the one with a hard expiry and a downstream configuration decision. Scope it to a decision memo, not a dashboard.
  5. Publish the allocation in one place with a one-line reason per item, so any escalation argues with the reasoning rather than with you, and the owners can see each other's deadlines.
  6. Hold back roughly one day. Something breaks most weeks, and an allocation with no slack fails in a way that damages all three commitments instead of one.
Follow-up
  • The acquisition owner says a rough number is better than nothing for a budget meeting. What exactly do you give them?
  • How would you decide whether the economy audit is genuinely urgent rather than merely important?
  • Two weeks of this pattern in a row. What structural change do you propose, and to whom?
  • 01

    Describe a time when you had to collaborate with a difficult team member.

  • 02

    A two-week live-ops event closed with gross bookings up 18% against the prior week. Your readout says it added no net revenue: summed price_usd_net in fct_iap_transaction over the 28 days spanning the event is flat against a matched window one release cycle earlier, and fct_currency_ledger shows premium-currency sinks falling while faucets from source_system = 'pass_tier' rose. The event owner disputes the readout in a review with their director present. Give the argument you would make in that room, and name the one piece of evidence you would concede.

  • 03

    Three requests land in the same week and you have five working days. The economy team wants a sink-to-faucet audit per currency after a faucet change shipped ten days ago. Live-ops wants a readout on an event that ends Friday, because the next event is configured from it. Acquisition wants 90-day net revenue per install by channel for a budget meeting in three weeks, and two of the channels launched six weeks ago. All three owners have escalated. Give your allocation, the reasoning you give each owner, and what you refuse or defer.

PracHub interview preparation framework ↗
Is this an official Epic Games interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Epic Games. Rounds and questions reflect what candidates have reported, not a process Epic Games has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How difficult is the interview process at Epic Games?

The interview process is considered rigorous, focusing on both technical skills and cultural fit. Candidates typically report needing several weeks of preparation to feel adequately ready.

PracHub interview research ↗
What differentiates successful candidates?

Successful candidates demonstrate strong technical expertise, effective problem-solving skills, and a genuine passion for gaming. They can clearly communicate their insights and collaborate with diverse teams.

PracHub interview research ↗
What is the company culture like at Epic Games?

The culture at Epic Games emphasizes innovation, creativity, and collaboration. The company seeks individuals who are not only skilled but also align with its core values and mission.

PracHub interview research ↗
How long does the interview process usually take?

The timeline can vary, often taking several weeks from the initial screen to the final decision, especially around holiday breaks.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.