As a Data Scientist at [24]7.ai, you play a pivotal role in shaping conversational artificial intelligence and customer engagement solutions that serve millions of global users daily. Your work bridges complex mathematical modeling with production-grade product features, helping enterprises automate interactions, optimize intent recognition, and drive measurable business outcomes. This role requires you to transform unstructured conversational data and massive telemetry logs into intelligent, self-optimizing systems.
You will collaborate closely with product managers, software engineers, and domain experts to design, test, and deploy machine learning models and experimentation frameworks. Whether you are building predictive intent engines, analyzing customer journey bottlenecks, or designing rigorous A/B tests to measure feature impact, your insights directly influence product roadmaps. The scale and complexity of conversational data at [24]7.ai mean your models must balance predictive accuracy with low-latency execution.
Expect an environment that demands both deep technical rigor and product-oriented thinking. You will not only build models but also diagnose metric shifts, defend your statistical methodologies, and communicate complex findings to non-technical stakeholders. Success in this role requires intellectual curiosity, resilience through multi-stage technical loops, and a relentless focus on creating value for end users.
Recruiter Conversation
reportedMost candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.
What to demonstrate
- Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
- Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
- The substance of the questions you ask back, which an experienced screener reads as a level signal
How to prepare
- Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
- Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
- Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
Hiring Manager Discussion
reportedUnderneath the questions about your past work sits a resourcing question. Given four things worth doing and one of you, which gets done and what happens to the rest? Managers ask because that is the daily texture of the job, and because the answer shows whether you rank work by effort or by what it changes. The weak version sorts by personal interest or by whoever asked most insistently. The strong version ties each candidate piece of work to a decision somebody downstream is waiting on, and then names the one you would drop and who you would tell.
What to demonstrate
- Whether you rank work by the decision it unblocks or by how interesting the method is
- How you describe a request you declined, and whether you can say who you said it to
- Whether your sense of how long something takes survives one follow-up question about the messy part
- How you decide something is good enough to hand over unfinished
How to prepare
- Write out your current queue and, next to each item, the decision that stays stalled until it lands. Anything with no waiting decision becomes your example of work you would cut
- Rehearse turning down a plausible stakeholder request out loud, including the smaller alternative you offered instead
- Have one case where you shipped a rough answer early and one where you refused to, with the reason that separated them
Technical Rounds
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Final Evaluation
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
PracHub editorial advice for the preparation topics above.
Randomising an experiment at the user level when users share an account
Two problems fire at once. Colleagues in one workspace see each other's work and talk to each other, so a treated user changes the behaviour of a control user in the same account, which violates the no-interference assumption and biases the estimate toward zero. Separately, outcomes within an account are strongly correlated, so the effective sample size is roughly n / (1 + (m - 1) * rho) for m users per account and intra-class correlation rho, not n. With rho around 0.3 and twenty users per account that is a design effect near 6.7, meaning a user-level confidence interval is about two and a half times narrower than it should be and results cross significance thresholds on noise alone. Randomise the account and cluster the standard errors.
Reporting a mean over accounts when account revenue is heavy-tailed
When a small number of accounts hold most of the revenue, the sample mean is dominated by whichever of them happens to be in the sample, and the sample variance keeps growing as more data arrives instead of stabilising. In that regime the usual central-limit-based confidence interval understates uncertainty, and a single renewal or a single large account's batch job can flip the sign of a measured effect. The fixes are to pre-register a winsorisation or capping rule before looking at the outcome, to report account counts crossing a threshold alongside the revenue figure, or to define the estimand on a bounded transform. Choosing the cap after seeing the result is a separate and worse problem, because the cap then encodes the answer.
Solving silently instead of narrating the reasoning
Say which branch you are taking and why you chose it over the alternative, for example checking the denominator first because it changes what the comparison means. A correct answer that arrives with no visible path scores below a rigorous one that needed a hint.
Dropping rows with missing values without naming the mechanism
Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the underlying mathematics behind hypothesis testing and how y…
Explain the underlying mathematics behind hypothesis testing and how you choose between parametric and non-parametric tests.
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Write down the assumption the method needs before you use the method.
- Translate the result into the decision it informs, in one plain sentence.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
What is the difference between frequentist and Bayesian approaches whe…
What is the difference between frequentist and Bayesian approaches when updating model priors in an online learning system?
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Bootstrap a confidence interval for net revenue retention
You have one row per account with arr_start_cents (ARR twelve months ago) and arr_end_cents (ARR today, zero if churned), covering the fixed cohort of accounts that had ARR twelve months ago. Net revenue retention is sum(arr_end_cents) / sum(arr_start_cents). Write a nonparametric bootstrap from scratch, without scipy.stats.bootstrap: resample accounts with replacement, recompute the ratio of sums on each resample, and return the point estimate with a 95% percentile interval from 10,000 resamples. Also report the interval you would get from the mean of per-account ratios, and explain the difference.
Approach
- Resample the account, because the account is the unit the estimand is defined over. One bootstrap draw is a vector of account indices and both numerator and denominator are recomputed from that same draw; resampling the two sides independently destroys the within-account correlation that makes a ratio estimator stable.
- Vectorise the draws: idx = rng.integers(0, n, size=(B, n)), then end[idx].sum(axis=1) / start[idx].sum(axis=1). A 10,000 by n index matrix is usually far cheaper than a Python loop; if the matrix is too large for memory, chunk over B rather than reverting to a loop.
- Take the interval from np.quantile(ratios, [0.025, 0.975]). The percentile interval differs from estimate +/- 1.96 * bootstrap SE whenever the resample distribution is skewed, which it will be here, and the skew is the thing you want represented.
- Compute the mean-of-ratios version on the same resamples, and state the exact relationship rather than guessing which of the two is larger. With r_i = arr_end_i / arr_start_i, the ratio of sums is the arr_start-weighted mean of exactly those r_i, so sum(end)/sum(start) - mean(r) = Cov(arr_start, r) / mean(arr_start) using the population covariance. The gap is positive when larger accounts retain and expand better than smaller ones, and negative when they do not; a cohort whose small accounts churn at a higher rate has positive covariance, which puts the mean of per-account ratios BELOW the ratio of sums. Requires arr_start_i > 0 for every account, which the fixed-cohort definition guarantees; r_i is floored at 0 and unbounded above, so a handful of 4x expansions among small accounts can flip the sign. Compute the covariance and report it instead of asserting a direction.
- Report the interval width beside the concentration of the cohort. If the largest account is 12% of starting ARR, a narrow interval is evidence that the resampling unit is wrong rather than evidence that the estimate is precise.
Worked solution 30 min
- start = df.arr_start_cents.to_numpy(float); end = df.arr_end_cents.to_numpy(float); n = len(start); point = end.sum() / start.sum()
- rng = np.random.default_rng(7); idx = rng.integers(0, n, size=(10_000, n)); ratios = end[idx].sum(1) / start[idx].sum(1)
- lo, hi = np.quantile(ratios, [0.025, 0.975]); return point, lo, hi
- per_acct = end / start; mean_point = per_acct.mean(); mean_boot = per_acct[idx].mean(1); compare np.quantile(mean_boot, [0.025, 0.975]) against (lo, hi), and report np.cov(start, per_acct, ddof=0)[0,1] / start.mean() as the quantity that accounts for the gap between the two centres.
Follow-up
- The cohort has 800 accounts and the largest is 12% of starting ARR. How much do you trust a percentile interval here?
- How would you extend this to an interval on the year-over-year change in NRR?
- Two accounts merged mid-window and one contract was co-termed into the other. How do you keep the cohort fixed?
Given a table of chat interactions, write a query using ranking functi…
Given a table of chat interactions, write a query using ranking functions to find the second-to-last customer touchpoint before escalation.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
How would you optimize a slow-running query that joins massive transac…
How would you optimize a slow-running query that joins massive transaction tables with real-time intent classification logs?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Net revenue retention on a cohort frozen twelve months back
fct_subscription_period carries subscription_period_id, account_id, arr_cents, plan_code, term_start_date, term_end_date, booked_at, amendment_type, superseded_by_id (the subscription_period_id of the version that replaced this one, null on the live version of a lineage) and is_current. Compute net revenue retention for month M: the summed arr_cents at M for the set of accounts holding arr_cents > 0 at M-12, divided by that same set's arr_cents at M-12. An account can hold more than one live subscription, a churned account contributes zero rather than dropping out, and nothing signed after M-12 may enter either side. Return the ratio plus the expansion, contraction and churn components in cents.
Approach
- Write one reusable as-of ARR snapshot parameterised by a date: rows whose term brackets the date AND whose booked_at is at or before the date, then only the version of each lineage that is still live at that date, then sum arr_cents per account. The booked_at guard matters because an amendment signed in advance otherwise co-exists with the term it replaces and double counts the account.
- Collapse the lineage on superseded_by_id, not on any attribute of the contract. Keep a row when superseded_by_id IS NULL, or when the successor it points at was booked after the date. Ranking with ROW_NUMBER() OVER (PARTITION BY account_id, plan_code ...) instead is wrong in both directions: an amendment that moves the account from one plan_code to another puts the old and new versions in different partitions, so both are rank 1, both bracket the date, and the account's arr_cents is counted twice; and two genuinely concurrent subscriptions that happen to share a plan_code land in one partition, so one of them is deleted.
- Resolve the successor with a LEFT JOIN back to fct_subscription_period on subscription_period_id, and treat a missing successor as not superseded. An inner join would silently delete an account's ARR on a dangling pointer, which is a data-quality bug in the source, not a retention movement.
- Never use is_current for the M-12 side. is_current describes today; using it at the historical snapshot backdates the present contract onto last year's cohort and makes retention look like 100 percent by construction.
- Freeze the cohort from the M-12 snapshot where arr_cents > 0, then LEFT JOIN the M snapshot onto it and COALESCE the missing side to zero. An inner join deletes exactly the churned accounts, which is the single largest way this number gets overstated.
- Return a ratio of sums, not a mean of per-account ratios. The two are different estimands: contraction is floored at zero while expansion is unbounded, so the mean of ratios is both biased relative to the aggregate and far noisier on a skewed revenue base.
- Decompose per account on the delta: positive delta is expansion, negative delta with a non-zero M value is contraction, a zero M value is churn. The three components must reconcile to numerator minus denominator.
- Prove no leakage: any account whose first contract began after M-12 must be absent from both sides, and the cohort row count must be identical in the numerator and denominator.
Worked solution 45 min
- Write arr_asof(d) as a CTE or lateral: from fct_subscription_period s take rows with term_start_date <= d AND term_end_date >= d AND booked_at <= d, LEFT JOIN fct_subscription_period succ ON succ.subscription_period_id = s.superseded_by_id, keep the row when s.superseded_by_id IS NULL OR succ.subscription_period_id IS NULL OR succ.booked_at > d, then sum arr_cents per account_id across every surviving version.
- Sanity-check the lineage rule on one amended account before going further: at a date after the amendment, the account must contribute exactly one version per lineage even when the amendment changed plan_code, term dates or both.
- Materialise base = arr_asof(M-12) filtered to arr_cents > 0, and curr = arr_asof(M).
- LEFT JOIN curr onto base on account_id and COALESCE(curr.arr_cents, 0) AS arr_now.
- Compute nrr = sum(arr_now)::numeric / NULLIF(sum(base.arr_cents), 0), and the three components with SUM(...) FILTER on the sign of arr_now - base.arr_cents and on arr_now = 0.
- Reconcile: assert sum(arr_now) - sum(base.arr_cents) = expansion - contraction - churn, and assert the cohort account count is identical on both sides.
Follow-up
- Net revenue retention can rise while the business shrinks. Show one mechanism and name the guardrail that catches it.
- How do you handle an account that co-terms two subscriptions into one mid-window, so the subscription count changes but the money does not?
- Finance computes this from invoiced amounts and gets a different number. Which is right for which question?
If daily active engagement drops by fifteen percent following a UI upd…
If daily active engagement drops by fifteen percent following a UI update, how would you systematically diagnose the root cause?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How would you approach feature selection when dealing with hundreds of…
How would you approach feature selection when dealing with hundreds of correlated variables in a customer interaction dataset?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
What metrics do you prioritize when evaluating an imbalanced classific…
What metrics do you prioritize when evaluating an imbalanced classification model for fraud or customer churn?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How would you design a metric to measure the success of an automated i…
How would you design a metric to measure the success of an automated intent-recognition bot in a customer service workflow?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
Describe how you would set up an experiment when network effects or in…
Describe how you would set up an experiment when network effects or interference between treatment and control groups are present.
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
How do you determine the required sample size and minimum detectable e…
How do you determine the required sample size and minimum detectable effect for an A/B test on a low-traffic support portal?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
How do you balance automated resolution rates with customer escalation…
How do you balance automated resolution rates with customer escalation rates when designing product KPIs?
Approach
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Separate novelty from durable lift in a console redesign
A redesigned request console was tested for four weeks, randomised at the account level. The lift in weekly successful interactive requests per account was +9.2%, +5.1%, +1.8% and +0.4% in experiment weeks one through four, and the week-four 95% interval spans plus or minus 3.1 percentage points. Accounts entered on a rolling basis as they next signed in, so week one is a different calendar week for different accounts. Say what the decaying series can and cannot establish, specify the readout that would settle it, and state what you would ship on Monday.
Approach
- Name the distinct explanations the same curve is consistent with, because the shipping decision differs across them: a novelty effect decaying to zero; a novelty effect decaying to a small positive plateau the test has no power to see; and a genuine effect whose measured size shrinks as control accounts learn about the change from colleagues or release notes.
- Fix the time axis before interpreting anything. With rolling entry, calendar week and weeks-since-first-exposure are different variables, and pooling them mixes the decay curve with a change in cohort composition, since the accounts that signed in on day one are systematically the most engaged. Recut on weeks-since-first-exposure and verify that plan_tier and pre-period usage mix are stable across entry cohorts.
- Read the interval width honestly. Plus or minus 3.1 percentage points at week four cannot separate a durable +1% from zero. The defensible statement is that no durable effect larger than roughly 3.5% was detected, not that there is no durable effect, and those two sentences lead to different decisions.
- Separate novelty from primacy using account tenure. Accounts created after launch have no prior console to be surprised by, so they cannot show novelty; if the redesign is genuinely better their curve should be flat or rising, while established accounts show the spike and decay. If both cohorts decay to zero, the effect really was novelty.
- Specify the settling readout rather than arguing about the four weeks you have. Hold back 5% of accounts as a long-run holdout for 90 days, and compare whole ISO weeks only: usage in this domain follows a hard five-to-two weekday cycle, and a window containing four business days instead of five moves this metric by several percent with no product change.
- Give the Monday answer. Ship if guardrails are clean and maintenance cost is low, because a decayed-to-zero effect with no harm is a neutral trade; but book none of the +9.2% in any forecast, and do not run a follow-up test on the same accounts inside the novelty window, because their baseline has not returned to steady state.
Worked solution 30 min
- Recut the four weekly estimates on weeks-since-first-exposure and confirm entry-week cohorts are comparable on plan_tier and pre-period usage.
- Split each week's estimate by account tenure, created before versus after launch, and compare the shapes of the two curves.
- Compute the horizon needed to halve the week-four interval: precision scales with the square root of exposure, so about four times the account-weeks are required.
- Write the ship note with the 5% holdout design, the 90-day re-read date, and an explicit statement of the effect sizes that remain unexcluded.
Follow-up
- What sample or horizon would you need to rule out a durable +1.5% at 80% power, given the week-four interval you have?
- How would you distinguish a novelty effect from a control arm that gradually learned about the change?
- The metric that decides renewal is up to eleven months away on annual contracts. What proxy do you use in the meantime, and how do you validate it once renewals land?
Monthly logo churn triples with no change in satisfaction
Monthly logo churn, computed as churned accounts divided by all paying accounts, tripled last month. Contracts are annual. From fct_subscription_period (account_id, term_start_date, term_end_date, amendment_type, auto_renew, booked_at, superseded_by_id, is_current) and dim_account (account_id, churned_at, account_status, employee_band, acquisition_channel), rebuild churn on the renewal-eligible base, separate calendar effects from customer behaviour, and state whether retention actually changed. Note that churned_at is sometimes set when the record was updated rather than at term end.
Approach
- Rebuild the denominator as accounts whose term_end_date falls in the month. An account eleven months from renewal sits in the current denominator while being structurally incapable of entering the numerator, so on annual contracts the published rate understates the truth by roughly the reciprocal of the annual renewal fraction and moves with the signing calendar.
- Plot the renewal-eligible base by month across two years. A signing surge twelve months earlier reproduces itself as an eligibility surge now, and a rate whose denominator ignores that tracks the sales calendar rather than customer sentiment.
- Date each churn by term_end_date, never by churned_at or an update timestamp. Inspect the distribution of churned_at minus term_end_date: a backlog cleared in one batch appears as a mass at a single date and shifts losses into whichever month the operations team did its paperwork.
- Apply the 45-day grace for late renewal paperwork so the most recent month and a half is marked not reportable, rather than printing a number that will rise once the paperwork lands.
- Compare corrected gross logo retention against its own trailing distribution, and if a real change survives, cut it by employee_band, acquisition_channel and plan_tier before proposing any cause.
Follow-up
- Some contracts in the window have not reached their renewal date. When does this require a survival estimator rather than a simple rate, and which one would you use?
- How do you report churn to an audience that wants a monthly number when the underlying event is annual and lumpy?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.
Describe a time when you disagreed with a cross-functional peer regard…
Describe a time when you disagreed with a cross-functional peer regarding a technical design or metric definition. How did you reach alignment?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Pick a story where you drove the decision, not one where you observed it.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Allocate one analyst week across three competing requests
Three requests arrive the same morning and you have one week. Finance wants per-account gross margin from fct_usage_daily for a pricing review in three weeks. Sales wants a renewal-risk list for accounts with term_end_date inside 60 days. A product manager wants an experiment readout for a decision being taken on Thursday. Produce your allocation with hours attached, what you say to whoever receives less, and one thing you refuse to do this week, with the reason each decision is defensible to the person it costs.
Approach
- The interviewer is probing whether you prioritise on decision timing and reversibility or on who asked most forcefully. Sort by the date each decision is actually taken and by what the default outcome is if nothing arrives.
- Apply that sort concretely. The Thursday readout has a hard irreversible deadline and no value afterwards. The pricing review has three weeks of slack. The renewal list has a rolling deadline set by term_end_date, so part of it is urgent this week and the rest is not, which means it can be split rather than deferred whole.
- Find the cheapest sufficient version of each request rather than the full version. The readout goes in full. The renewal list ships as a filtered query over renewal-eligible accounts ranked by two inspectable signals rather than as a model. The margin work is scoped to the accounts that dominate the pricing decision, since revenue is heavily skewed and the tail will not change the conclusion.
- Make the trade visible in one written note to all three at once, with dates. Telling each person separately that they are the priority is how an allocation becomes a credibility problem.
- Refuse something explicitly and say why. The model version of the renewal list is the usual candidate, because it cannot be evaluated without a holdout nobody has agreed to yet, and building it this week forecloses that.
- Leave slack. A plan with none is a plan to miss the one deadline that cannot move.
Follow-up
- The sales leader escalates to your manager. What did you already do that makes that a short conversation?
- Which of the three deadlines would you push back on, and what exactly would you ask for?
- What would you change about how these requests reach you so next week is not the same?
Report an underpowered consumption test to a non-technical executive
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Approach
- The interviewer is probing whether you can be decision-useful without either hiding the uncertainty or hiding behind it. Start from the decision rather than the statistics: establish what the executive would do differently at plus 4 percent versus zero, because if the action is identical the interval does not matter.
- Translate the interval into consequences in units the executive already reasons about. Multiply both endpoints by the cohort's baseline consumption and contracted rates to give an annualised revenue range, so the answer is a range of dollars rather than a range of percentages.
- Price the option to wait. Using the observed variance, state roughly how many additional account-weeks halve the interval width, so keep running becomes a quantified choice instead of a stall.
- Offer a cheaper path to the same decision: a lower-variance proximate outcome such as successful billable units on the new SKU, or CUPED using each account's pre-period consumption, quoting the expected variance reduction as one minus the squared pre-post correlation.
- Give a recommendation and name the single observation that would reverse it. A strong answer commits; a generic one recites the interval and leaves the decision on the table.
Follow-up
- The executive says it clearly works and is just not provable, so ship it. What is your answer?
- How much of the interval width comes from clustering and how much from the revenue tail, and what would you do about each?
- If you had to ship this week with no more data, which guardrail would you watch for the first fortnight and at what threshold would you roll back?
- 01
Describe a time when you disagreed with a cross-functional peer regarding a technical design or metric definition. How did you reach alignment?
- 02
Three requests arrive the same morning and you have one week. Finance wants per-account gross margin from fct_usage_daily for a pricing review in three weeks. Sales wants a renewal-risk list for accounts with term_end_date inside 60 days. A product manager wants an experiment readout for a decision being taken on Thursday. Produce your allocation with hours attached, what you say to whoever receives less, and one thing you refuse to do this week, with the reason each decision is defensible to the person it costs.
- 03
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Is this an official [24]7.ai interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at [24]7.ai. Rounds and questions reflect what candidates have reported, not a process [24]7.ai has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process for a Data Scientist at [24]7.ai?
The interview loop is considered difficult and rigorous, featuring multiple elimination rounds that test both theoretical mathematics and practical coding. Expect each peer-led round to challenge your assumptions and probe deeply into your foundational knowledge.
PracHub interview research ↗How should I prepare for unexpected mathematical questions?
Refresh your core college-level probability, statistics, and linear algebra. Practice solving problems manually and talking through your derivations out loud, as interviewers often present novel scenarios that require live mathematical problem-solving.
PracHub interview research ↗Is coding required in every round?
Not every round involves live coding, but a strong coding background is essential. Some rounds focus exclusively on mathematical concepts and resume projects, while others test your ability to write efficient queries and algorithmic solutions.
PracHub interview research ↗What is the best way to handle ambiguous questions during the interview?
Start by clarifying constraints, stating your assumptions clearly, and proposing a structured framework before diving into details. Interviewers appreciate candidates who pause to structure their thoughts rather than rushing into an answer.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22