A Data Scientist at Roku plays a critical role in shaping the user experience and driving business growth for millions of active streaming devices worldwide. Operating at the intersection of product, engineering, and business strategy, data scientists at Roku translate massive datasets into actionable insights that directly influence content discovery, user engagement, and platform monetization. Whether optimizing the home screen layout or refining ad-targeting algorithms, your work will have a tangible impact on how users interact with digital media.
Within the organization, data science teams are often embedded or aligned with specific product areas, such as the Personalization & Recommendations group. In these specialized roles, you will design, deploy, and analyze machine learning models and statistical frameworks that power Roku's content recommendation engines. Because the platform operates at an immense scale, the technical challenges are highly complex, requiring a sophisticated balance of data engineering efficiency, statistical rigor, and sharp product intuition.
Ultimately, Roku looks for data scientists who do not just build models in isolation but can act as strategic partners to product managers and software engineers. Successful candidates are those who can seamlessly transition from writing complex SQL queries and designing rigorous A/B tests to presenting high-level strategic recommendations to senior leadership.
Recruiter Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Screening
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Deep-Dive Interviews
reportedRounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.
What to demonstrate
- Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
- The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
- Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options
How to prepare
- Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
- Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
- Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
Conversations with Stakeholders
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
2 candidate reports. Individual accounts describe a particular role and hiring cycle.
Roku Backend Engineer Interview Experience — Rejected After Struggling With Multithreading and an Ad Exchange System Design
View report detailsRoku Data Scientist Interview Experience — HM Screen, SQL Round, Four Onsites, and an Offer
View report detailsPracHub editorial advice for the preparation topics above.
Testing hours, revenue or completion with a difference in means on a heavy-tailed distribution.
Listening and viewing hours per account are strongly right-skewed and content popularity is close to power-law, so the variance of a sample mean is dominated by a few accounts and the central limit approximation converges slowly at realistic sample sizes. A t-test on mean hours can flip sign when one heavy account's week changes, and an experiment can appear significant because a single title released into one arm's window. Capping at a pre-registered percentile, or decomposing into a rate (did they stream at all) and a conditional intensity, controls the variance, at the stated cost that capping biases toward zero exactly when the true effect lives in the tail.
Treating the account as the person, or the profile as the person.
A household account carries several people, profiles frequently are not switched, and shared-screen, car and speaker playback often lands on a default profile with no user behind it. Personalisation trained on a profile therefore learns a mixture, retention regressions attribute one member's behaviour to another, and a per-account taste statistic describes a household composite. The practical consequence is that apparent personalisation wins can be device or context effects, so any identity-level claim needs a stated unit and an acknowledgement of what that unit actually aggregates.
Dropping rows with missing values without naming the mechanism
Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.
Reading a dozen metrics with no multiplicity control
Nominate one primary metric before launch and treat the rest as guardrails or exploratory, with Bonferroni or Benjamini-Hochberg applied when you intend to make claims from them. Twenty independent tests at 0.05 under the null produce at least one false positive about 64 percent of the time.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
What are the core assumptions of Ordinary Least Squares (OLS) regressi…
What are the core assumptions of Ordinary Least Squares (OLS) regression, and how do you test for them?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Five invariant checks over a playback fact table
fct_stream arrives with stream_id, started_at, ended_at (nullable), played_seconds, max_position_seconds, completion_ratio, is_qualified, event_date, and duration_seconds joined from the content dimension (null for live events). Write one function returning a tidy frame with a row per rule: rule name, eligible rows, violations, violation share, one example stream_id. The five rules are duplicate stream_id; played_seconds exceeding wall-clock ended_at minus started_at; max_position_seconds exceeding duration_seconds; completion_ratio disagreeing with max_position_seconds divided by duration_seconds; and event_date not equal to the UTC date of started_at.
Approach
- Give every rule its own eligible-row mask before its violation mask. Rules 2 to 4 are undefined where ended_at or duration_seconds is null, and a comparison against NaN evaluates False, so a mostly-null column otherwise reports a clean bill of health.
- Compare the recomputed completion_ratio with a tolerance (absolute difference above 1e-6), never with ==. It is stored as a float division and exact equality fails on rows that are correct.
- Allow a small tolerance on rules 2 and 3 too: a final segment can carry the playhead a second or two past duration_seconds, and heartbeat timestamps come off client clocks. State the tolerance you chose rather than burying it.
- Return violations and eligible rows as separate columns so the share has a stated denominator, and carry one example stream_id per rule so the output is actionable rather than a number.
- Read rule 5 as a signal rather than corruption: event_date drifting from the UTC date of started_at is the signature of offline playback uploaded after the partition closed, and its distinct dates are the recompute list.
Follow-up
- Rule 5 fires on 0.4 percent of rows, all from one app_version, all with reported_at days after started_at. Is that a bug, and what do you do about the daily numbers already published?
- How would you turn these into a blocking check in the pipeline without failing the load every time one client version misbehaves?
- Which of the five would you expect to fire on live events specifically, and how do you keep them out of the denominator?
Simulate concurrent-stream refusals on a shared family account
A family account has four profiles and max_concurrent_streams = 2. Across an evening window of six hours, each profile independently attempts playback as a Poisson process at 0.75 attempts per hour. Durations are lognormal with a median of 24 minutes and sigma 0.6 on the log scale. An attempt arriving while two streams are already active is refused and abandoned, not retried or queued. Estimate the share of attempts refused and the mean refusals per account-evening, each with a 95 percent interval, then repeat with the cap at three.
Approach
- Simulate event-driven rather than on a time grid: draw exponential inter-arrivals per profile at rate 0.75 per hour, pool and sort the arrival times, and carry a short list of active end times.
- At each arrival, discard end times at or before that instant, then refuse if two remain. Only a served attempt draws a duration and pushes arrival plus duration; drawing a duration for a refused attempt and letting it hold a slot models a queue instead of a refusal.
- Start the window empty and do not warm it up, because the question is about an evening that begins with nobody streaming. Say explicitly that this makes the simulated share read below the steady-state value.
- Size the replications from a pilot: run 2,000 evenings, read the observed share, then solve 1.96 times the standard error to at most 0.002 using the between-evening standard deviation of the per-evening refusal share.
- Re-run with the cap at three by changing one constant and report the pair, because the decision this feeds is whether the cap is what the refusals are about.
Worked solution 30 min
- Write one evening as a function: per profile, cumulative exponential draws with scale 1/0.75 hours truncated at 6 hours, then concatenate and sort the arrival times.
- Sweep arrivals in order against a list of active end times: pop the expired, refuse if the list is already at the cap, otherwise draw a lognormal duration with mean parameter log(24/60) and sigma 0.6 in hours and append arrival plus duration.
- Return (attempts, refusals) per evening; run a 2,000-evening pilot, take the per-evening refusal share, and size the full run from its standard deviation.
- Run the sized number of evenings at cap 2 and again at cap 3, with a fixed seed per configuration.
- Report the pooled share (total refusals over total attempts) with a clustered interval built from the per-evening shares, plus the mean and 90th percentile of refusals per evening.
Follow-up
- What changes if a refused attempt retries after two minutes instead of being abandoned?
- Erlang-B gives a closed form for this. Where does it agree with your simulation and where does it not?
- How would you find these refusals in fct_stream, given that a refused attempt produces no stream row at all?
How would you structure a complex query using Common Table Expressions…
How would you structure a complex query using Common Table Expressions (CTEs) to aggregate monthly active users who stream specific genres?
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Explain how you would optimize a query that is running slowly due to l…
Explain how you would optimize a query that is running slowly due to large table joins on non-indexed columns.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Write a query utilizing a self-join and window functions to identify u…
Write a query utilizing a self-join and window functions to identify user engagement patterns across consecutive days.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Net revenue per active account-month without fanning out periods
fct_subscription_period has account_id, period_start_ts, period_end_ts, net_amount_usd and payment_status. fct_stream has account_id, started_at, played_seconds and is_qualified. Compute net revenue per active account-month for March 2026: the numerator is subscription revenue over periods with payment_status in ('paid','retried_paid'), recognised pro rata by the days of each period falling inside March; the denominator is distinct accounts with at least one qualified stream in March. A colleague's draft joins the two tables on account_id, filters streams to qualified March rows, and sums net_amount_usd; its numerator is roughly forty times too large. Explain the factor and write the correct query.
Approach
- Name the mechanism and decompose the number. The join is one period to many streams, so each period row is duplicated once per matching stream and its net_amount_usd is added that many times: the naive sum equals the un-prorated period sum weighted by each account's March qualified-stream count, which is the net_amount_usd-weighted mean stream count over the accounts the inner join keeps. The distance from there to the correct numerator has two further parts, in opposite directions — pro-rating shrinks the correct figure, while the inner join has dropped every period belonging to an account that streamed nothing in March. A weighted mean near thirty streams against a pro-rata factor near 0.75 is what puts the draft around forty times high; the number is not arbitrary, which is why it looks plausible enough to ship.
- Adopt the rule that produces the fix: two fact tables never meet at row grain. Aggregate each to account grain first and join the aggregates, or — better here — do not join at all, because the numerator and the denominator are independent scalars that share no key.
- State the interval convention before writing the arithmetic. Treating periods as half-open [period_start_ts, period_end_ts) makes consecutive periods tile without overlap; overlap_days = GREATEST(0, LEAST(period_end_ts, DATE '2026-04-01') - GREATEST(period_start_ts, DATE '2026-03-01')) in days, and the contribution is net_amount_usd * overlap_days / total_days_of_period. An annual period contributes roughly 31/365 of its amount.
- Build the denominator on its own: COUNT(DISTINCT account_id) FROM fct_stream WHERE is_qualified AND started_at >= '2026-03-01' AND started_at < '2026-04-01'. Do not restrict it to paying accounts — the metric deliberately holds subscription and advertising revenue to one engaged-account denominator, and filtering to payers is what makes a tier migration look like two unrelated numbers moving.
- State the closing rule with the result: the month must be at least 45 days closed, because refunds and chargebacks are recorded against periods late and an open month always reads high.
Worked solution 35 min
- Reproduce the bug deliberately: compute the naive joined sum, and beside it the un-prorated SUM(net_amount_usd) over the same period rows restricted to accounts with at least one qualified March stream. Their ratio is the fan-out multiple, and it is the only ratio in this problem that has a closed form.
- Write the overlap expression and verify that one period's contributions, summed across every month it touches, reconstruct its net_amount_usd exactly.
- Compute the denominator independently and combine the two as one-row CTEs joined with CROSS JOIN, so neither side can fan out the other.
- Re-run for February and March and confirm no revenue is counted in both, which is the boundary test for the half-open convention.
Follow-up
- One account holds two overlapping periods after a plan migration. Whose revenue counts for March, and does your pro-rata still sum to each period's own total?
- Advertising revenue belongs in this numerator. What grain does it arrive at, and what breaks in your query when you add it?
- How would you turn this into a monthly time series without rewriting the date literals, and what would you do about periods that span three months?
What metrics would you use to define the success of a newly redesigned…
What metrics would you use to define the success of a newly redesigned Roku home screen layout?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Imagine you launch an A/B test for a new recommendation algorithm. The…
Imagine you launch an A/B test for a new recommendation algorithm. The primary engagement metric goes up, but ad-clicks go down. How do you decide whether to roll out the feature?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Describe how you would design an experiment to test a new push notific…
Describe how you would design an experiment to test a new push notification campaign, including sample size calculation and minimum detectable effect (MDE).
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
What is the difference between correlation and causation, and what sta…
What is the difference between correlation and causation, and what statistical methods can you use to establish the latter in non-experimental data?
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Publish a per-person metric when no table holds a person
Leadership asks for streams per user per week on a weekly dashboard. You have dim_account (account_id, plan_tier, max_profiles), dim_profile (profile_id, account_id, profile_type, personalisation_opt_out, last_active_ts) and fct_stream (profile_id, account_id, device_type, started_at, played_seconds, is_qualified). Neither table holds a person. Deliverable: state the unit you will publish, define the metric fully, name what that unit actually aggregates, and give one diagnostic that tells the dashboard's readers how much of the data has no identifiable person behind it.
Approach
- Establish that both available units are wrong in different directions: an account is a household so it aggregates several people, and a profile is not reliably one person either because profiles are frequently not switched and shared-screen, car and speaker playback lands on whatever profile was last active or on a default.
- Pick the account as the published unit and justify it by what the number will be used for: billing, churn and revenue denominators are all per account, so an account-grained engagement metric joins to every other number on the dashboard without an impedance mismatch, while a profile-grained one silently changes the denominator between panels.
- Define it fully: qualified hours per active account-week, numerator sum(played_seconds)/3600 over is_qualified rows, denominator distinct (account_id, ISO week) pairs with at least one qualified stream, is_test_account excluded, window the four most recent complete ISO weeks, with the account-level two-distinct-date rule used for the active-account count.
- Label what the unit aggregates directly on the dashboard, not in documentation: the figure is hours per subscribing household per week, and family and duo plan_tier accounts read higher for composition reasons and are not more engaged people.
- Give the diagnostic that makes the ambiguity visible: share of qualified hours on profile_type = 'default_unnamed' plus the share on device_type in ('smart_speaker','car','smart_tv'), published beside the headline, since that is the fraction of consumption whose person is genuinely unknown and it tells a reader how much interpretive weight the number bears.
- Rule out the false fix explicitly: never divide account hours by max_profiles, because max_profiles is a plan entitlement that does not vary with household size and dividing by it manufactures a per-person figure that is a deterministic function of the plan.
Worked solution 25 min
- Write one sentence each for what dim_account and dim_profile actually identify, and state the two mechanisms that break the profile-equals-person assumption.
- Write the chosen metric in full, including numerator, denominator, window, exclusions and the two-distinct-date active rule, and state the reason the account grain matches the other dashboard panels.
- Write the unidentified-hours diagnostic as a query shape: share of sum(played_seconds) where profile_type = 'default_unnamed' or device_type in ('smart_speaker','car','smart_tv'), over total qualified played_seconds, by week.
- Write the label that sits under the headline number naming the household interpretation, in the words a non-analyst reader would need.
- State the prohibition on max_profiles division and the one-line reason, so the next person to ask for per-person numbers gets an answer rather than a repeat.
Follow-up
- Personalisation claims a win measured at the profile level. What do you need to know before you believe it is a person-level effect and not a device or context effect?
- A cut of the dashboard by plan_tier shows family accounts at three times individual accounts on hours. Write the sentence you would put under that chart.
- If you could add one column to dim_profile to make this less ambiguous, what would it be and what would it still fail to tell you?
Slate click rate jumped overnight with no ranker change
Home-row click rate — was_clicked over rendered rows in fct_impression (impression_id, slate_id, surface, slate_position, rendered_at, viewport_visible_ms, ranker_version, was_clicked) — rose from 4.1% to 6.8% on a single day and held there. ranker_version is unchanged across the step. Streams attributed to home_row in fct_stream (impression_id, start_source, started_at) did not rise. Using impression and stream data only, find the cause and state what you tell the ranking team, who have already claimed the improvement in a review deck. Deliverable: root cause plus the correctly restated series.
Approach
- Ask which side of the ratio moved. A rate that steps while the downstream count is flat is almost always a denominator event, so plot clicks and impressions as separate absolute series before touching the ratio again. Do the arithmetic first as well: with the numerator fixed, a rate going from 4.1% to 6.8% pins the denominator to 4.1/6.8 of its former size, so you know the size of the impression loss to look for before you open the first query.
- If impressions fell, locate the fall: group by surface, slate_position band and viewport_visible_ms = 0. Loss concentrated in high positions and in never-visible rows is the signature of a client or collector dropping below-the-fold renders, which removes mostly non-clicked rows and lifts the rate arithmetically with no behaviour change.
- Prove rows are missing rather than never rendered. Left-join fct_stream rows with start_source = 'algorithmic_slate' and non-null impression_id back to fct_impression and count the ones that no longer resolve. Orphaned attributions date the breakage to the hour and are hard to argue with.
- Look at impressions per slate_id, and require it to reconcile with the rate step. A twenty-item row now logging twelve is a truncation bug in the renderer or collector, and it accounts for the loss exactly if the distinct slate_id count is flat; a stable per-slate count with fewer slates is a sampling or partition drop upstream. The two have different owners and different backfill stories, and any per-slate figure that implies a larger loss than the rate step requires means a second thing is also wrong.
- Restate the series on a basis the defect does not touch — clicks per slate_id, or the rate restricted to slate_position <= 8 where logging is intact — and take the dated evidence to the ranking team before the number reaches the review, not after.
Follow-up
- Impressions in the truncated tail are missing for six days. Can an experiment that ran in that window still be read, and under exactly what assumption?
- What monitor catches this within an hour, and what does it cost in false alarms on a normal release day?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.
How do you handle selection bias when designing an observational study…
How do you handle selection bias when designing an observational study rather than a randomized controlled trial?
Approach
- Pick a story where you drove the decision, not one where you observed it.
- Close with what you would do differently, concretely.
- Quantify the outcome, including what you would not claim credit for.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that project again?
Walk through an analysis you shipped that was wrong
Describe a case where you delivered a result that was later shown to be wrong, and it had already been acted on. Cover how the error surfaced, whether you or someone else found it, what the wrong number caused, and what you changed afterwards. Prepare an example whose root cause was a definition, a denominator or a join — not a transcription slip. The interviewer will push on the mechanism, not the apology. Deliverable: a four-minute account that ends with a specific control now running in a pipeline.
Approach
- The probe is whether your account has a mechanism in it. Choose an error that generalises — a denominator that silently changed population, a join that fanned rows, a metric partitioned on event_date while offline playback arrived days late and landed in the wrong partition — rather than one that only teaches you to check your typing.
- State the blast radius factually and early: which decision was taken, how long the number stood, what it cost. A candidate who softens this is answering a different and easier question, and the interviewer can hear the substitution.
- Explain how it surfaced without adjusting who found it. The generalisable detail is why your own checks did not catch it, which is a statement about your checks rather than about your luck.
- Name the control you added and where it now lives: a row-count assertion after the fan-out join, a reconciliation that recomputes a closed day after late-arriving offline playback and alerts above a threshold, a denominator assertion inside the query. A fix that lives in a pipeline is different in kind from a resolution to be more careful.
- Close with whether the control has fired since, or how you tested that it would. That single sentence is what separates a fix from an intention, and interviewers ask for it when candidates do not offer it.
Follow-up
- Why didn't your own review catch it? Be specific about what you did check.
- What class of error would that control still not catch, and what would you add next?
- Have you found an error in someone else's published analysis since? How did you raise it?
Disagree with a product manager about a completion metric
A product manager proposes making duration-normalised completion rate the team's primary metric for the quarter, and wants a plain global unweighted version "so it's simple enough to put on a wall." You expect the unweighted version to move several points on catalogue mix alone, with no change in how satisfying anything was. You have one meeting, they own the roadmap, and you will work with them for years. Deliverable: the argument, the evidence you bring, and the fallback you accept if they still want the simple version.
Approach
- The probe is whether you can lose the decision without damaging either the metric or the relationship. Bring the failure already reproduced: recompute the unweighted global rate over the last two quarters and point at the weeks it moved several points where the only input that changed was which content people played.
- Decompose the variance instead of objecting. Attribute the movement in the unweighted rate to between-cell mix — (content_type, duration decile) — versus within-cell movement. A decomposition is an argument; an assertion that mix matters is a preference.
- Concede the real cost honestly: the mix-weighted version is harder to explain and harder to recompute. Bring the mitigation rather than dismissing the objection — weights fixed from a stated reference month, published once, so anyone can reproduce the number without redoing the weighting.
- Offer the fallback deliberately and in advance: ship both, with the unweighted rate labelled a diagnostic and the weighted one as the decision metric, and pre-agree the one observation that would mean the simple version has misled the team.
- Leave the decision with them, with the consequence written down before it happens. That is what makes a later correction a shared prediction coming true rather than a retrospective argument about who was right.
Follow-up
- They ship the simple version and it moves four points during a heavy release week. How do you raise it without saying you told them so?
- Is there a case where the unweighted rate is genuinely the right primary metric?
- Suppose you are wrong and the mix effect turns out to be small. What would you change about how you argued this?
- 01
How do you handle selection bias when designing an observational study rather than a randomized controlled trial?
- 02
Describe a case where you delivered a result that was later shown to be wrong, and it had already been acted on. Cover how the error surfaced, whether you or someone else found it, what the wrong number caused, and what you changed afterwards. Prepare an example whose root cause was a definition, a denominator or a join — not a transcription slip. The interviewer will push on the mechanism, not the apology. Deliverable: a four-minute account that ends with a specific control now running in a pipeline.
- 03
A product manager proposes making duration-normalised completion rate the team's primary metric for the quarter, and wants a plain global unweighted version "so it's simple enough to put on a wall." You expect the unweighted version to move several points on catalogue mix alone, with no change in how satisfying anything was. You have one meeting, they own the roadmap, and you will work with them for years. Deliverable: the argument, the evidence you bring, and the fallback you accept if they still want the simple version.
Is this an official Roku interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Roku. Rounds and questions reflect what candidates have reported, not a process Roku has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the recruiter screening at Roku?
Candidates are often surprised by the technical nature of the initial recruiter screening. Be prepared for the recruiter to ask high-level technical questions regarding your experience with SQL, Python, and statistical modeling, in addition to standard behavioral questions.
PracHub interview research ↗What is the coding environment like during the SQL technical interview?
The technical screening is typically conducted via HackerRank or a similar live-coding platform. You will be expected to write clean, executable SQL queries to solve problems of varying difficulty. Depending on the interviewer, you may not be allowed to run exploratory queries, so writing structured, accurate code on your first attempt is highly valued.
PracHub interview research ↗How does Roku evaluate culture fit?
Roku values candidates who are collaborative, self-motivated, and highly focused on delivering measurable business impact. They look for professionals who can navigate ambiguity, communicate clearly with non-technical partners, and take extreme ownership of their analytical projects.
PracHub interview research ↗What is the typical timeline from the first screen to an offer?
The timeline can vary significantly. While some candidates complete the process in approximately three weeks, others report a longer timeline of six to eight weeks, particularly if the role involves multiple specialized rounds or stakeholder interviews across different time zones.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22