As a Data Scientist at Jerry.ai, you sit at the intersection of consumer finance, insurance, and personalized AI. Jerry.ai is on a mission to simplify the way people manage their financial lives, and your role is to extract actionable intelligence from vast amounts of user data to drive product innovation. You will be responsible for building models that optimize pricing, improve customer retention, and personalize the user experience across Jerry.ai's core service offerings.
This position is critical to the company’s ability to scale efficiently. You will work closely with product and engineering teams to translate business ambiguity into rigorous technical solutions. Whether you are improving A/B testing frameworks or developing predictive models for insurance underwriting, your work directly influences the company's bottom line and the day-to-day financial health of its users. Expect a fast-paced environment where your ability to bridge the gap between complex data and clear business strategy is highly valued.
Recruiter Screen
reportedWhoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.
What to demonstrate
- Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
- Whether your reason for wanting the role points at the work itself rather than the company's reputation
- Whether your language signals the level being screened for: what you decided yourself versus what you were handed
How to prepare
- Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
- Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
- Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
Take-Home Assignment
reportedThe clock is part of the test. Three to six hours is not enough to do everything the dataset supports, so the submission mostly reveals how you spend a fixed budget against an open question. A reviewer sees which paths you took and, by absence, which you abandoned. Work that runs out of time inside the analysis ships a thin conclusion, while work that cuts scope early protects the last hour for writing. The most reliable way to lose here is to leave the scoping decision implicit, so it reads as something you missed rather than something you chose.
What to demonstrate
- Whether the scope you settled on is presented as a decision with a reason, rather than left for the reader to infer from what is missing
- Whether the depth of the work is consistent with the stated time budget, instead of several half-finished directions left open
- Whether the closing section reads as something written on purpose rather than assembled from whichever cells survived
How to prepare
- Run a timed rehearsal on a public dataset with a hard stop, holding the final sixty minutes for writing no matter where the analysis has got to
- Before opening the data, list the questions it could plausibly answer, pick one, and keep the discarded ones as a short note on what you did not attempt and why
- Commit a one-line finding after each analysis step so the writeup is assembled from recorded results rather than from memory at midnight
Case Study Interview
reportedUnderneath the business framing, this round is usually asking whether you can turn a fuzzy goal into a quantity that could be computed from data such a business would plausibly hold. That means a metric with a stated numerator, denominator, eligibility rule and time window, plus an honest account of the conditions under which it would mislead you. Answers come apart when a candidate names a familiar metric and never defines it, because every follow-up then lands on an ambiguity that was left open and the candidate has to invent the definition under pressure.
What to demonstrate
- Whether a named metric arrives with its denominator, eligibility rule and window attached rather than assumed
- Whether the measure follows from the mechanism you proposed, or is a recognisable metric retrofitted to it afterwards
- Whether you name a guardrail that would reveal the gain came from somewhere you did not want it to come from
- Whether you can say what data the plan requires and what you would settle for if that logging were never implemented
How to prepare
- Take five metrics you reach for by reflex and write each as one sentence containing numerator, denominator, eligibility rule and time window. The ones you cannot finish are the ones that will fail under follow-up.
- For a product you use daily, write the measurement plan you would propose for a change to it: primary metric, one guardrail, the unit of analysis, and the table the numbers would come from.
- Practise the substitution question. For three metrics you like, write what you would measure instead if the event you depend on were not being logged.
1 candidate reports. Individual accounts describe a particular role and hiring cycle.
Jerry.Ai Software Engineer Interview Experience — GPS Trivia in the Recruiter Screen, Stuck on "move" in the OA
Recruiter Screen The interviewer was really friendly. She started by walking me through the company, then asked if I had any questions for her. In the middle she threw in an unusual question: how GPS works / the underlying design. Roughly, how GPS works, how the signal propagates, and why errors happen — for example the difference between the satellite's atomic clock and the clock on your phone.…
Read full experiencePracHub editorial advice for the preparation topics above.
Comparing accounts that received a sales or customer-success touch against those that did not
Assignment of coverage is deliberate and pulls in both directions at once: the largest accounts get a named owner because they are valuable, and the accounts showing distress get one because they are at risk. The comparison therefore mixes a strong positive selection with a strong negative one, and the naive estimate can come out with either sign depending on which assignment rule dominated during the period examined. Nothing about matching on observed size fixes this, because the risk signal that triggered coverage is usually the same signal that predicts the outcome. It needs either an actual randomised or staggered rollout of coverage, or a design built on a capacity constraint or territory boundary that assigns coverage for reasons unrelated to account health.
Randomising an experiment at the user level when users share an account
Two problems fire at once. Colleagues in one workspace see each other's work and talk to each other, so a treated user changes the behaviour of a control user in the same account, which violates the no-interference assumption and biases the estimate toward zero. Separately, outcomes within an account are strongly correlated, so the effective sample size is roughly n / (1 + (m - 1) * rho) for m users per account and intra-class correlation rho, not n. With rho around 0.3 and twenty users per account that is a design effect near 6.7, meaning a user-level confidence interval is about two and a half times narrower than it should be and results cross significance thresholds on noise alone. Randomise the account and cluster the standard errors.
SQL that silently fans out on a one-to-many join
State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.
Answering a product-sense question with a list of features
Answer with a decision and the measurement that would settle it: the hypothesis, the primary metric, the guardrails, and the result that would make you not ship. A feature brainstorm cannot be wrong, which is exactly why it earns no points.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How would you approach building a churn prediction model for our insur…
How would you approach building a churn prediction model for our insurance products?
Approach
- Set a baseline first, so any model has something honest to beat.
- Say how the offline result would be validated online before it is trusted.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Audit daily usage rows for grain and arithmetic violations
You are handed fct_usage_daily as a pandas DataFrame with account_id, workspace_id, sku_code, usage_date, billable_quantity, included_quantity_applied, overage_quantity, list_amount_cents, discount_amount_cents, net_amount_cents, cogs_cents, is_restated, first_written_at and restated_at. The declared grain is one row per (account_id, workspace_id, sku_code, usage_date). Write audit(df) returning a DataFrame with one row per failing check: check name, failing row count, and one example key. Cover at minimum grain duplication, negative quantities or amounts, the identity net = list - discount, billable = included + overage, and rows where is_restated is true but restated_at is null.
Approach
- Check the grain before anything else with df.duplicated(subset=key, keep=False), and count rows rather than groups so a key appearing twice contributes 2 — if the grain is broken every arithmetic count below it is uninterpretable.
- Express each invariant as a boolean Series over the whole frame. The cent columns are integers and compare exactly, so use !=; the numeric(18,6) quantity columns need np.isclose with atol=1e-6 because included + overage is a decimal sum.
- Handle null as its own failure mode. Comparisons against NaN return False, so a check written as rows_that_pass = (a == b - c) silently files every null-amount row wherever the negation happens to land; build each check as violations = ~condition | column.isna().
- Collect the checks as a list of (name, mask) pairs and assemble the output in one pass, so adding a check is one line and every check reports in the same shape.
- Order the output with structural failures (grain, null keys) above arithmetic failures, and report zero-count checks too — a check that silently disappears when it passes is indistinguishable from a check that was never run.
Follow-up
- Which of these should block a dashboard refresh and which should only warn?
- Rows with is_restated = true legitimately change value after first write. How do you make yesterday's audit result reproducible?
- How would you extend this to catch a partition that is missing entirely rather than wrong?
Bootstrap a confidence interval for net revenue retention
You have one row per account with arr_start_cents (ARR twelve months ago) and arr_end_cents (ARR today, zero if churned), covering the fixed cohort of accounts that had ARR twelve months ago. Net revenue retention is sum(arr_end_cents) / sum(arr_start_cents). Write a nonparametric bootstrap from scratch, without scipy.stats.bootstrap: resample accounts with replacement, recompute the ratio of sums on each resample, and return the point estimate with a 95% percentile interval from 10,000 resamples. Also report the interval you would get from the mean of per-account ratios, and explain the difference.
Approach
- Resample the account, because the account is the unit the estimand is defined over. One bootstrap draw is a vector of account indices and both numerator and denominator are recomputed from that same draw; resampling the two sides independently destroys the within-account correlation that makes a ratio estimator stable.
- Vectorise the draws: idx = rng.integers(0, n, size=(B, n)), then end[idx].sum(axis=1) / start[idx].sum(axis=1). A 10,000 by n index matrix is usually far cheaper than a Python loop; if the matrix is too large for memory, chunk over B rather than reverting to a loop.
- Take the interval from np.quantile(ratios, [0.025, 0.975]). The percentile interval differs from estimate +/- 1.96 * bootstrap SE whenever the resample distribution is skewed, which it will be here, and the skew is the thing you want represented.
- Compute the mean-of-ratios version on the same resamples, and state the exact relationship rather than guessing which of the two is larger. With r_i = arr_end_i / arr_start_i, the ratio of sums is the arr_start-weighted mean of exactly those r_i, so sum(end)/sum(start) - mean(r) = Cov(arr_start, r) / mean(arr_start) using the population covariance. The gap is positive when larger accounts retain and expand better than smaller ones, and negative when they do not; a cohort whose small accounts churn at a higher rate has positive covariance, which puts the mean of per-account ratios BELOW the ratio of sums. Requires arr_start_i > 0 for every account, which the fixed-cohort definition guarantees; r_i is floored at 0 and unbounded above, so a handful of 4x expansions among small accounts can flip the sign. Compute the covariance and report it instead of asserting a direction.
- Report the interval width beside the concentration of the cohort. If the largest account is 12% of starting ARR, a narrow interval is evidence that the resampling unit is wrong rather than evidence that the estimate is precise.
Worked solution 30 min
- start = df.arr_start_cents.to_numpy(float); end = df.arr_end_cents.to_numpy(float); n = len(start); point = end.sum() / start.sum()
- rng = np.random.default_rng(7); idx = rng.integers(0, n, size=(10_000, n)); ratios = end[idx].sum(1) / start[idx].sum(1)
- lo, hi = np.quantile(ratios, [0.025, 0.975]); return point, lo, hi
- per_acct = end / start; mean_point = per_acct.mean(); mean_boot = per_acct[idx].mean(1); compare np.quantile(mean_boot, [0.025, 0.975]) against (lo, hi), and report np.cov(start, per_acct, ddof=0)[0,1] / start.mean() as the quantity that accounts for the gap between the two centres.
Follow-up
- The cohort has 800 accounts and the largest is 12% of starting ARR. How much do you trust a percentile interval here?
- How would you extend this to an interval on the year-over-year change in NRR?
- Two accounts merged mid-window and one contract was co-termed into the other. How do you keep the cohort fixed?
Explain how you would optimize a slow-running SQL query.
Explain how you would optimize a slow-running SQL query.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
How would you join multiple tables to identify high-value customer seg…
How would you join multiple tables to identify high-value customer segments?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Net revenue retention on a cohort frozen twelve months back
fct_subscription_period carries subscription_period_id, account_id, arr_cents, plan_code, term_start_date, term_end_date, booked_at, amendment_type, superseded_by_id (the subscription_period_id of the version that replaced this one, null on the live version of a lineage) and is_current. Compute net revenue retention for month M: the summed arr_cents at M for the set of accounts holding arr_cents > 0 at M-12, divided by that same set's arr_cents at M-12. An account can hold more than one live subscription, a churned account contributes zero rather than dropping out, and nothing signed after M-12 may enter either side. Return the ratio plus the expansion, contraction and churn components in cents.
Approach
- Write one reusable as-of ARR snapshot parameterised by a date: rows whose term brackets the date AND whose booked_at is at or before the date, then only the version of each lineage that is still live at that date, then sum arr_cents per account. The booked_at guard matters because an amendment signed in advance otherwise co-exists with the term it replaces and double counts the account.
- Collapse the lineage on superseded_by_id, not on any attribute of the contract. Keep a row when superseded_by_id IS NULL, or when the successor it points at was booked after the date. Ranking with ROW_NUMBER() OVER (PARTITION BY account_id, plan_code ...) instead is wrong in both directions: an amendment that moves the account from one plan_code to another puts the old and new versions in different partitions, so both are rank 1, both bracket the date, and the account's arr_cents is counted twice; and two genuinely concurrent subscriptions that happen to share a plan_code land in one partition, so one of them is deleted.
- Resolve the successor with a LEFT JOIN back to fct_subscription_period on subscription_period_id, and treat a missing successor as not superseded. An inner join would silently delete an account's ARR on a dangling pointer, which is a data-quality bug in the source, not a retention movement.
- Never use is_current for the M-12 side. is_current describes today; using it at the historical snapshot backdates the present contract onto last year's cohort and makes retention look like 100 percent by construction.
- Freeze the cohort from the M-12 snapshot where arr_cents > 0, then LEFT JOIN the M snapshot onto it and COALESCE the missing side to zero. An inner join deletes exactly the churned accounts, which is the single largest way this number gets overstated.
- Return a ratio of sums, not a mean of per-account ratios. The two are different estimands: contraction is floored at zero while expansion is unbounded, so the mean of ratios is both biased relative to the aggregate and far noisier on a skewed revenue base.
- Decompose per account on the delta: positive delta is expansion, negative delta with a non-zero M value is contraction, a zero M value is churn. The three components must reconcile to numerator minus denominator.
- Prove no leakage: any account whose first contract began after M-12 must be absent from both sides, and the cohort row count must be identical in the numerator and denominator.
Worked solution 45 min
- Write arr_asof(d) as a CTE or lateral: from fct_subscription_period s take rows with term_start_date <= d AND term_end_date >= d AND booked_at <= d, LEFT JOIN fct_subscription_period succ ON succ.subscription_period_id = s.superseded_by_id, keep the row when s.superseded_by_id IS NULL OR succ.subscription_period_id IS NULL OR succ.booked_at > d, then sum arr_cents per account_id across every surviving version.
- Sanity-check the lineage rule on one amended account before going further: at a date after the amendment, the account must contribute exactly one version per lineage even when the amendment changed plan_code, term dates or both.
- Materialise base = arr_asof(M-12) filtered to arr_cents > 0, and curr = arr_asof(M).
- LEFT JOIN curr onto base on account_id and COALESCE(curr.arr_cents, 0) AS arr_now.
- Compute nrr = sum(arr_now)::numeric / NULLIF(sum(base.arr_cents), 0), and the three components with SUM(...) FILTER on the sign of arr_now - base.arr_cents and on arr_now = 0.
- Reconcile: assert sum(arr_now) - sum(base.arr_cents) = expansion - contraction - churn, and assert the cohort account count is identical on both sides.
Follow-up
- Net revenue retention can rise while the business shrinks. Show one mechanism and name the guardrail that catches it.
- How do you handle an account that co-terms two subscriptions into one mid-window, so the subscription count changes but the money does not?
- Finance computes this from invoiced amounts and gets a different number. Which is right for which question?
How would you measure the success of a new recommendation engine on th…
How would you measure the success of a new recommendation engine on the Jerry.ai app?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you decide between two different pricing strategies using hi…
How would you decide between two different pricing strategies using historical data?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Propose a metric to track user engagement and explain why it is better…
Propose a metric to track user engagement and explain why it is better than current alternatives.
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
If our conversion rate drops by 5% overnight, how would you investigat…
If our conversion rate drops by 5% overnight, how would you investigate the cause?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Explain the process of designing and evaluating an A/B test for a new …
Explain the process of designing and evaluating an A/B test for a new feature.
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
What statistical methods do you use to determine if a trend is signifi…
What statistical methods do you use to determine if a trend is significant?
Approach
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Build an activation metric tree for a self-serve onboarding rebuild
The self-serve signup flow is being rebuilt. You have dim_account (account_id, created_at, is_internal, first_workspace_at, is_current) and fct_api_request (account_id, workspace_id, environment, api_key_id, http_status, traffic_class, request_at). Build a three-layer metric tree from signup through to production dependency, name the primary metric for this specific change, and state the fixed measurement window and the denominator for each layer. Two of the three layers share a denominator and the third does not, so say which invariant still holds across the tree once the denominators differ. Say which layer the onboarding team can actually move and which it cannot, and why raw signup counts are not on the tree.
Approach
- Fix the cohort once: accounts created in cohort period P with is_internal = false, collapsed to one row per account_id since dim_account is type 2. Say out loud that employee, demo and load-test accounts are removed, otherwise the tree partly measures internal testing.
- Layer one, qualified signup, on the cohort denominator: accounts recording at least one successful request (http_status < 400) with traffic_class in ('interactive','batch') within 14 days of created_at. This replaces raw signup counts, which measure registration-form friction and duplicate sign-ups from one organisation rather than demand.
- Layer two, activation, also on the cohort denominator: the share of the cohort whose first successful request with a non-null api_key_id arrives within a fixed 168 hours of created_at. Fixed window, not trailing, so that cohort weeks are directly comparable and no cohort is censored.
- Layer three, production depth, on a conditional denominator: the share of activated accounts holding a workspace with environment = 'production' whose trailing 7-day count of successful non-synthetic requests exceeds a floor calibrated from the data, for example the volume above which 90-day retention stops rising. A floor picked by intuition makes this a metric about the floor. Conditioning on activation is the right operational choice, because the question it answers is whether accounts that got started go on to depend on the product, but it means this rate is not on the same scale as the two above it: a conditional rate of 0.50 sits perfectly happily underneath a 0.30 activation rate. To place it back on the funnel, multiply it by the activation rate and report that product as the cohort share.
- Align the traffic_class filter across all three layers, otherwise the underlying account sets do not nest: an activation definition that admits traffic_class = 'ci' will count accounts that layer one excluded, and the tree stops being a funnel. With mixed denominators the invariant that survives is on account sets and therefore on counts, not on rates. Then name the primary metric for this change, the seven-day activation rate, because the onboarding flow controls the path to first successful call and nothing below it. Production depth is the lagging check that the activation gain was real rather than one sample request.
Worked solution 25 min
- Build the cohort table: one row per account_id, bucketed by the ISO week of created_at, is_internal = false, dim_account collapsed to is_current = true.
- For each layer, find the first qualifying request per account from fct_api_request under that layer's filter and keep min(request_at), using EXISTS or a lateral so the cohort stays at one row per account.
- Compute layers one and two as shares of the cohort denominator and layer three as a share of the activated accounts, carrying the raw account count beside every rate, and mark any cohort week not yet clear of the longest window (14 days) as not reportable.
- Plot the three rates by cohort week, and separately confirm nesting on the counts within each cohort.
Follow-up
- Activation rises 6 points and the production-depth rate among activated accounts is flat 60 days later. Say what happened to the cohort share of production-depth accounts, and why a flat conditional rate is the better of the two readings available here.
- Median time-to-first-successful-call answers the same question with more information. Why is it a worse weekly dashboard metric, and what estimator would make it honest?
Activation drops six points starting at a deploy hour
Seven-day activation, defined as an account's first request with http_status < 400, api_key_id not null and traffic_class <> 'synthetic_monitor' within 168 hours of created_at, fell six points for sign-up cohorts after a Tuesday. A client SDK major version shipped that morning. From fct_api_request (account_id, api_key_id, sdk_name, sdk_version, http_status, traffic_class, request_at) and dim_account (account_id, created_at, is_internal), decide whether activation actually fell or the metric's inputs changed, and state in advance what evidence would convince you of each.
Approach
- Decompose the definition and recompute activation under each relaxation: status only, status plus traffic_class, then the full definition. If the entire drop lives in the api_key_id clause, this is an instrumentation question rather than a behavioural one.
- Measure the null rate of api_key_id by sdk_version and by hour. A stamping change appears as a step at the deploy boundary confined to the new version; a behavioural change appears as a ramp that grows with adoption and leaves old-version traffic untouched.
- Hold the cohort's SDK mix fixed before comparing. New sign-ups adopt the newest version first, so a cohort-level drop can be pure composition even when no individual version changed at all.
- Corroborate with a source the release did not touch: whether the same cohorts appear in fct_usage_daily with non-zero billable_quantity, and whether their fct_support_ticket rows with category = 'onboarding' rose.
- Write the decision rule down before looking at the answer. An instrumentation artefact predicts unchanged downstream usage and a version-confined null step; a real regression predicts falling downstream usage and more onboarding tickets in the same cohorts.
Follow-up
- Old-version and new-version populations are not exchangeable, because new accounts adopt the new version first. How would you build a comparison that is not confounded by cohort age?
- What backfill or metric-versioning policy keeps the historical series interpretable once you fix the stamping?
Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and set a baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
- Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
- Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.
Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02One query pattern, written three times
- Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
- On the third attempt, write the grain of every CTE as a comment before writing its body.
- Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.
Deliverable: Three independent versions of the same query plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Only the statistics you will be asked to defend
- Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
- Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
- Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.
Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04One case, and the assumptions holding it up
- Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
- Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
- Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.
Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
- Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in a longer weekend block
- Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
- Re-read only your own notes from this week, and open no new material.
- Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Nearly every data role forces a trade between the analysis you want and the one that fits the decision window. Prepare a case where you deliberately shipped something less rigorous, named the weakness to the person relying on it, and said what would change your answer. The naming is the part interviewers listen for.
How do you handle missing or noisy data in a production pipeline?
How do you handle missing or noisy data in a production pipeline?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Quantify the outcome, including what you would not claim credit for.
- Close with what you would do differently, concretely.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Announce a metric fix that cuts the headline number
Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.
Approach
- The interviewer is probing whether you can land a correction as an operational change with a plan attached, rather than as an announcement other people then have to clean up after.
- Quantify each exclusion separately before telling anyone: internal accounts, synthetic monitors, load tests. Three known quantities are a discussion; one alarming total is an argument.
- Be precise about which side of the metric each exclusion touches, because one team's goal is on a count and the other's is on a ratio. The traffic-class filters remove requests, so they shrink the numerator only. Dropping internal accounts removes them from the ratio's denominator as well, since internal accounts carry ordinary account_status values and therefore sit in that denominator. Internal accounts are active in almost every week while the real base is not, so the numerator loses a larger share than the denominator and the ratio falls by less than the count does. Compute both and say which one the 11 percent is before anybody assumes.
- Check whether the trend changes, not only the level. A constant 11 percent shift is a rebasing and nothing more. A shift that widens over time means the reported growth was partly internal or synthetic, which makes the existing goals unachievable as written and changes what you are asking teams to do.
- Sequence the disclosure: the metric owner and the two teams whose goals move first and privately, then the board channel with a written bridge, then the dashboard. The dashboard is last because a number that changes without explanation is read as instability rather than as a fix.
- Run both series for one reporting period with the bridge visible, restate history rather than letting the series break at a date, and set the date the old series is removed.
- Propose the goal treatment yourself: rebase each target by the shift measured on the metric that target is written against, rather than leaving each team to negotiate individually, which is where corrections of this kind usually die.
Follow-up
- One team's quarterly goal is now unreachable. Rebase the target or let it miss, and what does each choice teach the organisation?
- How would this have been caught when the metric was first defined?
- What else on that dashboard shares this failure mode, and how would you find out this week?
Report an underpowered consumption test to a non-technical executive
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Approach
- The interviewer is probing whether you can be decision-useful without either hiding the uncertainty or hiding behind it. Start from the decision rather than the statistics: establish what the executive would do differently at plus 4 percent versus zero, because if the action is identical the interval does not matter.
- Translate the interval into consequences in units the executive already reasons about. Multiply both endpoints by the cohort's baseline consumption and contracted rates to give an annualised revenue range, so the answer is a range of dollars rather than a range of percentages.
- Price the option to wait. Using the observed variance, state roughly how many additional account-weeks halve the interval width, so keep running becomes a quantified choice instead of a stall.
- Offer a cheaper path to the same decision: a lower-variance proximate outcome such as successful billable units on the new SKU, or CUPED using each account's pre-period consumption, quoting the expected variance reduction as one minus the squared pre-post correlation.
- Give a recommendation and name the single observation that would reverse it. A strong answer commits; a generic one recites the interval and leaves the decision on the table.
Follow-up
- The executive says it clearly works and is just not provable, so ship it. What is your answer?
- How much of the interval width comes from clustering and how much from the revenue tail, and what would you do about each?
- If you had to ship this week with no more data, which guardrail would you watch for the first fortnight and at what threshold would you roll back?
- 01
How do you handle missing or noisy data in a production pipeline?
- 02
Weekly active organisations, the count on the company dashboard, has never excluded rows where dim_account.is_internal is true, and it counts traffic with traffic_class in synthetic_monitor and load_test. Correcting both reduces that count by 11 percent and removes most of the growth reported over two quarters. The figure appears in a board deck and in two teams' quarterly goals, one written on the count and one on the weekly active organisation ratio, whose denominator is accounts whose account_status was in ('trial','free','active_paid') through the week. Decide the order in which you tell people, what the dashboard shows during the transition, and what you propose happens to goals already set against the old definition.
- 03
An account-randomised packaging change ran six weeks across 900 paying accounts. The effect on billable units per account per month is plus 4.1 percent, with a 95 percent interval from minus 3.2 to plus 11.8 after clustering standard errors at the account and applying the pre-registered winsorisation at the 99th percentile. An executive with no statistical background wants one number this week to decide a full rollout. Produce a three-sentence spoken answer, one chart, and an explicit recommendation of ship, stop or keep running, with the cost of each option stated.
Is this an official Jerry.ai interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Jerry.ai. Rounds and questions reflect what candidates have reported, not a process Jerry.ai has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
The process can range from a few weeks to over a month, depending on the speed of the hiring team and the number of rounds. Stay proactive in your communication with the recruiter.
PracHub interview research ↗Are the take-home assignments reasonable?
While some candidates find them challenging, they are designed to test your real-world problem-solving skills. Prioritize clarity, documentation, and business reasoning over just finding the "correct" answer.
PracHub interview research ↗What is the company culture like?
Jerry.ai is a fast-paced, startup-oriented environment. Success here requires high autonomy, a bias for action, and the ability to thrive in ambiguous situations.
PracHub interview research ↗Should I ask for feedback?
Yes, always ask for feedback, though it is not guaranteed. Focus on what you can learn from each interaction to improve for future rounds.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22