A Data Scientist at Hopper sits at the intersection of product innovation, financial engineering, and predictive modeling. Hopper is not just a travel booking platform; it is a fintech powerhouse that leverages massive datasets to eliminate anxiety from travel planning. From predicting future flight prices to managing the risk profiles of products like Price Freeze and Cancel for Any Reason, data science is the engine that drives the company’s revenue and customer retention.
As a Data Scientist, you will be responsible for transforming billions of real-time search and pricing data points into actionable product features. You will collaborate closely with product managers, business leaders, and engineers to design algorithms that predict market volatility and optimize user conversion. Your work directly impacts how millions of travelers budget for their trips, making this role both highly visible and intellectually challenging.
The work environment at Hopper is fast-paced, highly autonomous, and deeply quantitative. To succeed here, you must possess not only strong technical skills in programming and statistics but also a keen product sense. The team values individuals who can look at a complex, messy dataset and extract strategic insights that can be immediately deployed to improve the user experience and drive business growth.
Recruiter Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Take-Home Data Challenge
reportedThe clock is part of the test. Three to six hours is not enough to do everything the dataset supports, so the submission mostly reveals how you spend a fixed budget against an open question. A reviewer sees which paths you took and, by absence, which you abandoned. Work that runs out of time inside the analysis ships a thin conclusion, while work that cuts scope early protects the last hour for writing. The most reliable way to lose here is to leave the scoping decision implicit, so it reads as something you missed rather than something you chose.
What to demonstrate
- Whether the scope you settled on is presented as a decision with a reason, rather than left for the reader to infer from what is missing
- Whether the depth of the work is consistent with the stated time budget, instead of several half-finished directions left open
- Whether the closing section reads as something written on purpose rather than assembled from whichever cells survived
How to prepare
- Run a timed rehearsal on a public dataset with a hard stop, holding the final sixty minutes for writing no matter where the analysis has got to
- Before opening the data, list the questions it could plausibly answer, pick one, and keep the discarded ones as a short note on what you did not attempt and why
- Commit a one-line finding after each analysis step so the writeup is assembled from recorded results rather than from memory at midnight
Solution Walkthrough
reportedAn added round often puts you in front of someone outside the core hiring team: a partner engineer, a product owner, a domain expert, sometimes a more senior manager. The question they are really asking is not whether you can do the work but whether they would trust a number that came from you. That changes what a good answer looks like. Lead with what the decision cost and what it changed, keep the method available but not central, and be plain about the limits of your evidence. Overstating a result is the fastest way to lose this round.
What to demonstrate
- Whether you can explain a technical choice to someone who will never read your code, without either flattening it into nothing or hiding inside jargon
- Honesty about evidence strength: what the analysis establishes, what it only suggests, and what it cannot say at all
- How you take disagreement, specifically whether you update on a good objection, hold your position with reasons, or fold on contact
How to prepare
- Write the two-sentence version of your most technical project for a non-specialist, then check that neither sentence needs a method name to make sense.
- For one result you are proud of, write the strongest objection someone could raise and a response that concedes the part of it that is correct.
- Prepare one decision that turned out to be wrong: how you found out, what it cost, and what you changed afterwards. A senior cross-functional interviewer asks for this more often than a technical one does.
Deep-Dive Interviews
reportedAn extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.
What to demonstrate
- Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
- Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
- How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one
How to prepare
- Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
- Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
- Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
PracHub editorial advice for the preparation topics above.
Using the search as the demand unit when the trip is the demand unit
One trip intent generates dozens of searches over days, across devices, mostly while logged out, and the number of searches per intent is a property of the interface rather than of demand. Any product change that encourages comparison or re-sorting inflates the denominator, so look-to-book falls while the product improves, and a change that reduces re-searching raises the metric while nothing about demand moved. The logged-out majority also means traveller_id is NULL for most early-funnel rows, so joining searches to bookings on traveller_id silently drops the part of the funnel you were trying to measure. Collapse on trip_intent_key first, then count, and report searches per intent separately as a diagnostic rather than letting it sit inside a conversion rate.
Reading cancellation, completion or repeat rates on cohorts that have not matured
A cohort of bookings made last week for stays six months out cannot have cancelled at the check-in gate yet, so its cancellation rate is mechanically near zero and its completion rate mechanically near zero as well, in opposite directions. Comparing that cohort with a mature one is not a noisy comparison, it is a guaranteed wrong one, and the bias always makes the recent period look different in a way that invites a false story about a recent change. Because lead time is heavily right-skewed, the mean lead time is a bad maturity threshold; use the cohort's 95th percentile, or report a hazard at a fixed age (cancelled within k days of booking) with k capped at the youngest cohort's elapsed age. The same applies to repeat rate, where the honest answer is often that the cohort in question is not readable for another nine months.
Accepting a metric definition without asking about the denominator
Pin down the denominator, the eligibility filter and the time window before computing anything: conversion rate per session, per user, per eligible user and per new user are four different numbers with different behaviour. Restate the definition in one sentence and get agreement before you analyse.
Comparing periods without accounting for seasonality or day-of-week
Compare whole weeks against whole weeks and check whether the same swing appeared in prior cycles or prior years before attributing it to anything you changed. Weekday and weekend populations often differ enough that a Tuesday-to-Saturday comparison is meaningless.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Solve this probability puzzle: If a traveler has a specific probabilit…
Solve this probability puzzle: If a traveler has a specific probability of booking a flight on day one, how does that probability compound over a seven-day window given changing prices?
Approach
- Sanity-check the answer against a simple bound or a simulated case.
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- What sample size would you need to detect an effect half this size?
- Which assumption here is most likely to be violated in practice?
What machine learning algorithms would you consider for a real-time re…
What machine learning algorithms would you consider for a real-time recommendation engine, and how would you evaluate their performance?
Approach
- Check what information would not exist at prediction time, and exclude it.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
How would you set up an A/B test to evaluate a new push notification a…
How would you set up an A/B test to evaluate a new push notification algorithm designed to encourage bookings?
Approach
- Say how the offline result would be validated online before it is trusted.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- How would you choose the decision threshold, and who owns that choice?
Permutation test occupancy on market-week randomised clusters
A ranking change was randomised over 48 clusters, each a destination market crossed with a stay week, 24 per arm. You get clusters with destination_market_id, stay_week, arm, stayed_nights, sellable_nights; markets recur across several weeks. The metric is occupancy computed as a ratio of sums, treatment minus control. Write a permutation test from scratch with 10,000 reshuffles and a bootstrap confidence interval, using no scipy or statsmodels testing function. Report the observed difference, a two-sided p-value and a 95% interval, and state the resampling unit you chose for each.
Approach
- Compute the observed statistic as a ratio of sums within each arm, not a mean of per-cluster occupancies, so a market-week with 9,000 sellable nights does not carry the same weight as one with 300.
- For the permutation null, shuffle the arm label vector across clusters. Randomisation was independent per market-week, so the cluster is the exchangeable unit; permuting the night-level rows instead destroys the cluster correlation and shrinks the null distribution by roughly the square root of the nights per cluster, which turns almost any observed difference into a significant one.
- Vectorise the reshuffles: generate a (10000, 48) matrix of random values, argsort each row, and use the first 24 positions as the treatment index set, so the whole null distribution is built with array operations rather than a Python loop over 10,000 iterations.
- Use the add-one p-value, (1 + count of |permuted| >= |observed|) / (B + 1). The uncorrected version can report exactly zero, which claims more certainty than 10,000 reshuffles can support.
- For the interval, resample whole markets with replacement rather than clusters, because a market's weeks share demand conditions and are not independent draws; then recompute the ratio-of-sums difference on each resample. Report the implied minimum detectable effect at 48 clusters alongside the p-value, so a null result is read as underpowered rather than as evidence of no effect.
Worked solution 35 min
- obs = stayed[arm=='t'].sum()/sellable[arm=='t'].sum() - stayed[arm=='c'].sum()/sellable[arm=='c'].sum().
- Build idx = rng.random((10_000, 48)).argsort(axis=1); treatment mask = first 24 columns; compute both arms' ratio-of-sums per row with matrix multiplication against the stayed and sellable vectors.
- p = (1 + (np.abs(perm_diffs) >= abs(obs)).sum()) / (10_000 + 1).
- Bootstrap: sample market ids with replacement, gather all their clusters, recompute the difference 10,000 times, take the 2.5th and 97.5th percentiles.
- Compute the MDE: 2.8 times the permutation null's standard deviation, and report it next to the p-value.
Follow-up
- Your interval contains zero at 48 clusters. What cluster count would you need for an 80% chance of detecting a 2 percentage point occupancy lift, and what does that cost in calendar time?
- Why is a per-traveller randomisation on this same change biased, and in which direction?
- Which guardrails would you read alongside occupancy, and which failure mode does each one catch?
Rank market-weeks by qualified intents that never converted
Using the qualifying-intent definition over fct_search (trip_intent_key, searched_at_utc, destination_market_id, check_in_date, results_returned_count, is_bot_flagged) and fct_booking (trip_intent_key, booked_at_utc), produce a supply-acquisition list: the 20 (destination_market_id, ISO week of check_in_date) pairs with the most qualifying intents that never converted within seven days of their first search. Note that fct_booking.trip_intent_key is NULL for deep-link, offline and partner-channel bookings. Return market, check-in week, non-converting intents and total intents.
Approach
- Reuse the qualifying-intent set: non-bot, results_returned_count > 0, collapsed to one row per trip_intent_key carrying its first search timestamp and the check_in_date from that first search.
- Write the anti-join as NOT EXISTS, correlated on trip_intent_key with the seven-day window inside the subquery. NOT EXISTS is NULL-safe and lets the planner choose an anti-join.
- If you insist on NOT IN, add WHERE trip_intent_key IS NOT NULL to the inner query. Without it, one NULL in the subquery makes the predicate UNKNOWN for every outer row and the result set is empty.
- Group by destination_market_id and DATE_TRUNC('week', check_in_date), counting both non-converting and total intents so the reader sees volume and share together.
- Order by non-converting intents descending and limit to 20. Ordering by share alone promotes market-weeks with five intents and one miss.
Follow-up
- Which of these market-weeks are a supply shortage and which are a restriction or price problem? Which column separates them, and what join do you need?
- You run the query and get zero rows. Walk me through how you find out why, without changing the query at random.
- How would you exclude intents that did convert under a different trip_intent_key, such as the same traveller shifting their dates by a week?
Twelve-month repeat-trip rate on mature first-stay cohorts
fct_booking has traveller_id, booking_id, check_in_date, check_out_date and booking_status with values 'confirmed', 'amended', 'cancelled', 'no_show', 'in_stay' and 'stayed'. Compute the twelve-month repeat-trip rate: for travellers whose first stayed booking checked out in a given calendar month, the share with a second stayed booking whose check_out_date falls within 365 days of the first. Report cohort month, cohort size, repeat travellers and the rate. Publish only cohorts that have had a full 365 days to mature.
Approach
- Restrict to booking_status = 'stayed' before anything else. 'confirmed' and 'in_stay' bookings have not been delivered, and 'amended' must be checked against the lifecycle before you assume it is live.
- Number each traveller's stays with ROW_NUMBER() OVER (PARTITION BY traveller_id ORDER BY check_out_date, booking_id). The booking_id tiebreak makes the result deterministic when two stays end on the same date, which happens whenever a party books two units.
- Pivot the first two stays with conditional aggregation, MIN(CASE WHEN rn = 1 THEN check_out_date END) and the same for rn = 2, or self-join rn = 1 to rn = 2. Keep the second stay even when it lands outside 365 days, so the denominator is unaffected by the numerator condition.
- Decide and state whether a second booking ending the same day as the first counts as a repeat trip; requiring the second stay's check_in_date to be after the first stay's check_out_date is the defensible default and excludes same-trip extra rooms.
- Cohort on DATE_TRUNC('month', first check_out_date) and suppress with an explicit filter any cohort whose month end is less than 365 days before the data as-of date.
Worked solution 30 min
- Count stayed bookings and distinct travellers holding at least one; the second figure bounds the cohort universe.
- Build the ranked stay table and verify that every traveller's rn = 1 row does carry their minimum check_out_date.
- Produce first and second stay pairs, then compute the within-365-day flag.
- Group by cohort month and drop immature cohorts with a WHERE clause rather than by trimming the chart afterwards.
- Compare cohorts twelve months apart rather than adjacent months, so seasonal position is held fixed.
Follow-up
- Your July cohort reads six points below January. Before calling that a regression, what do you check?
- How does the answer change if you cohort on first booking creation instead of first completed stay, and which question does each version answer?
- A traveller's first stay was fully refunded after check-out. Does it still open a cohort, and what does your choice do to the denominator?
How would you design a dashboard to track the performance and user ado…
How would you design a dashboard to track the performance and user adoption of the Price Freeze feature?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
What metrics would you look at to determine if a user should buy a fli…
What metrics would you look at to determine if a user should buy a flight fare immediately or wait for a price drop?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you measure the long-term customer lifetime value (LTV) of a…
How would you measure the long-term customer lifetime value (LTV) of a user who purchases a fintech product versus one who does not?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Explain how you would handle a scenario where your primary metric show…
Explain how you would handle a scenario where your primary metric shows positive movement, but guardrail metrics show a negative trend.
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you determine the appropriate sample size and duration for a hi…
How do you determine the appropriate sample size and duration for a high-traffic product experiment?
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- Decide the analysis before seeing data, including how long it runs and when you look.
- State the primary metric and the minimum effect worth shipping, then size the test.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Measure a loyalty threshold effect with regression discontinuity
Travellers reach the gold loyalty_tier in dim_traveller at 10 completed stay-nights within a rolling 12 months; the rule is deterministic and evaluated nightly. Leadership wants the causal effect of gold status on completed stay-nights in the following 90 days, and tier cannot be randomised. Specify the design, the running variable, the estimand you can actually identify, and the two validity checks that would kill it. Then say what you conclude if the density of travellers at exactly 10 nights is 2.3 times the density at 9.
Approach
- Set up a sharp RDD: assignment is a deterministic step function of completed nights in the rolling window, so identification comes from continuity of potential outcomes at the cutoff rather than from randomisation. Estimate with local linear regression on each side, a triangular kernel, a data-driven bandwidth, and bias-corrected robust confidence intervals rather than conventional ones.
- Name the estimand before quoting it: a local average treatment effect for travellers sitting at 9 to 10 nights. It says nothing about travellers at 3 nights or at 40, and the difference matters because the business question is usually about the latter.
- Handle the discreteness of the running variable. Nights are integers with most mass below 15, so the local window holds very few distinct values; use inference designed for a discrete running variable rather than treating it as continuous, and never cluster the standard errors on the running variable itself.
- Run the manipulation test (a density test at the cutoff) and read a 2.3x jump at 10 over 9 as a failed check on its face, since it is consistent with travellers adding a night to qualify, which is self-selection into treatment and breaks continuity.
- Distinguish deliberate manipulation from mechanical heaping before condemning the design: integer stay lengths pile at round totals whether or not a threshold exists, so compare the histogram against a period before the threshold existed or against a cohort on a different threshold. If the heaping predates the rule, it is mechanical; if it appeared with the rule, it is manipulation.
- Confirm nothing else changes at 10 nights. If a coupon, an email trigger or a cancellation-policy change fires at the same cutoff, the discontinuity identifies the bundle and not gold status, and that has to be stated in the headline rather than in a footnote.
Worked solution 45 min
- Define the running variable as completed nights in the rolling 12-month window at the nightly evaluation, and freeze the outcome window as the 90 days after that evaluation date.
- Fit local linear regressions either side with a triangular kernel and a data-driven bandwidth, and report the bias-corrected robust interval alongside the conventional one.
- Run the density test at the cutoff and record the 2.3x ratio, then compare the integer-nights histogram against a pre-threshold period to classify the heaping as mechanical or manipulated.
- Run placebo RDDs at 8 and at 12 nights and confirm both return estimates indistinguishable from zero.
- If the heaping is manipulation, re-estimate as a donut RDD excluding a symmetric window around the cutoff, report the precision cost, and state plainly if the design no longer identifies the effect.
Follow-up
- The threshold is evaluated nightly on a rolling 12-month window, so travellers cross and fall back repeatedly. What does that do to the design and to the running variable?
- The density test fails and manipulation is real. Name an instrument you would consider for tier status and state the exclusion restriction it needs.
- Leadership actually wants the effect at 40 nights. What design gives them that, and what would it cost?
Net revenue per sellable night fell six percent quarter over quarter
Net revenue per sellable night fell 6.1% quarter over quarter on the stay-date axis. Occupancy fell 4.6% and ADR fell 1.6%. Sellable nights rose 8%, and supply grew as planned. Using fct_stay_night, fct_rate_availability_snapshot and dim_supply_unit (property_class, listed_at_utc, listing_status, units_blocked), decompose the fall into rate and occupancy, then each of those into mix and within-segment, and determine how much is the mechanical effect of new supply entering the denominator before it sells. Deliverable: the decomposition with every term sized, and one recommendation.
Approach
- Verify the identity on your own row set before decomposing anything: revenue per sellable night equals ADR times occupancy exactly, because revenue/sellable = (revenue/stayed) times (stayed/sellable), and only when all three are computed over the same units and dates. Work in logs so the components add, and confirm that -1.6% and -4.6% reproduce -6.1%.
- Run shift-share on each component separately and on the same segmentation, market by property_class by lead-time bucket, using stayed nights as the weight base for ADR and sellable nights for occupancy. Each rate must be decomposed on the denominator it is actually built from, or the terms will not reconcile.
- Isolate the supply-cohort effect explicitly. Split units by listed_at_utc into those live in both quarters and those listed during the quarter, and recompute all three metrics on the stable cohort alone. New listings add sellable nights immediately and revenue only once they ramp, so the gap between the all-supply and stable-cohort figures is dilution, and dilution is not a performance problem.
- Check the denominator for the opposite artefact. A supplier blocking cheap dates raises this metric while shrinking supply, so confirm the sellable-night convention did not change between quarters, and report sellable nights as a level rather than only as a denominator.
- Assemble one table with within-segment rate, mix rate, within-segment occupancy, mix occupancy, new-supply dilution and residual, and either make the residual small enough to ignore or explain it.
- Recommend only after the table exists. If the fall is dilution plus mix, cutting price attacks the one term that was already healthy and makes the next quarter worse.
Follow-up
- New supply accounts for 70% of the fall. At what point does that stop being dilution and start being a ramp problem you should act on?
- How would you report this metric so that next quarter's supply growth does not make the number look good for the wrong reason?
- Which of these terms could a supplier-facing team actually move, and over what horizon?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.
Retract an occupancy comparison after the decision shipped
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
Approach
- Quantify the error before telling anyone, because the first question will be 'how wrong'. Recompute both markets on sellable nights, which excludes units_blocked, and report the corrected gap and its sign, not just that the original was wrong.
- Separate the numerical error from the decision error. The comparison was invalid, but the spend move may still have been right; establishing whether the decision would have flipped is a different analysis and it is the one the business needs.
- Tell the decision-maker directly and first, before it appears in a dashboard or a peer surfaces it. Order matters because being told by a third party converts a mistake into a credibility problem.
- Write the note with the correction, the size, the decision implication, and the reversal cost in that order. Three weeks of moved spend has a real cost to undo, and a correction that does not price the reversal forces the reader to do the work you skipped.
- Name the specific control that would have caught it and put it in place in the same note: an assertion in the query that both arms use the same denominator expression, and the denominator named in the chart title. A retraction without a mechanism reads as an apology rather than a fix.
Follow-up
- The corrected numbers still support the original decision. Do you still send the note, and does it read differently?
- Your manager suggests quietly fixing the dashboard and not raising it. How do you respond?
- What would you have had to do differently three weeks ago, in the query itself, rather than in your review habits?
Defend a cancellation finding against the team it damages
A supplier-growth team moved one market's inventory from 'strict' to 'flexible' cancellation_policy in fct_rate_availability_snapshot. Their dashboard shows bookings up 14% over eight weeks. Your read on fct_stay_night shows stayed nights flat, and traveller-initiated cancellation up from 18% to 31% in the same lead-time bucket. Their quarterly goal is booked nights, and the lead has already sent the 14% to their director. You have ten minutes in their weekly review. Prepare what you open with, what you concede, and what you will not soften.
Approach
- Before the meeting, rebuild both periods as a hazard at a fixed age: share cancelled within k days of booked_at_utc, with k capped at the elapsed age of the youngest cohort. The flexible-policy cohort is younger, so a raw cancellation rate would be low for maturity reasons alone, and presenting that comparison hands the room a correct objection that kills the finding on its first sentence.
- Open by conceding the part that is true and theirs: bookings did rise 14%, the campaign did what it was designed to do on the booking axis. Naming their win first removes the reading that you are attacking the team rather than the metric.
- State the disagreement as an axis disagreement, not a competence one: booked nights is counted on booked_at_utc, stayed nights on stay_date, and the gap between them is exactly what a cancellation-policy change moves. Show the two series on one chart with both axes labelled.
- Quantify the delivered outcome in their own units so the tradeoff is arithmetic rather than opinion: stayed nights flat means the incremental bookings cancelled at roughly the rate that absorbs the whole 14%, and contribution margin per stayed night absorbed the servicing and payment-processing cost of the bookings that did not convert.
- Offer a route that keeps their goal intact: propose booked-nights-net-of-cancellation as the team's tracked number, or a stay-date readout held until the cohort matures, and say which you would commit to defending upward on their behalf.
Follow-up
- The lead says your cancellation cohort is not mature enough to compare. Walk me through the exact calculation that makes it comparable.
- Their director asks you directly whether the campaign should be rolled back. What do you say, and what would change your answer?
- How would you have set this up eight weeks ago so this conversation never happened?
Disagree with a product manager about the denominator
A product manager wants to ship a filter-refinement panel that lets travellers re-sort and narrow results more easily. Early data shows it increases searches per trip_intent_key from 6.2 to 9.4. The team's headline metric is bookings divided by fct_search rows, which falls from 4.1% to 2.9% under the change. The PM reads this as the feature hurting conversion and wants to kill it. You think the denominator is wrong. Prepare how you make that case in a working session without winning on a technicality.
Approach
- Recompute the metric on the defensible denominator first and bring both numbers: DISTINCT trip_intent_key values with at least one booking within 7 days of the first search for that key, over DISTINCT keys with is_bot_flagged = FALSE and results_returned_count > 0. If intent-level conversion is flat or up while search-level conversion falls, the disagreement resolves itself in one table.
- Explain the mechanism rather than the rule: searches per intent is a property of the interface, so any feature that encourages comparison inflates the denominator and any feature that discourages it improves the metric while demand is unchanged. That makes the current metric reward a worse product, which is an argument the PM has reason to care about.
- Concede what the search count does tell you and keep it: report searches per intent as a separate diagnostic, since 6.2 to 9.4 may be healthy exploration or may be people failing to find anything, and those have opposite implications.
- Distinguish the two hypotheses with evidence rather than assertion: compare results_returned_count and time-to-first-booking per intent between arms. Rising refinement with stable intent conversion and stable zero-result rate reads as exploration; rising refinement with a rising zero-result rate reads as failure to find.
- Agree the decision rule with the PM before reading the result, so the metric change is not seen as moving the goalposts after the fact, and put intent-level conversion on the team's dashboard alongside the old number rather than replacing it silently.
Follow-up
- Intent-level conversion is also down, by 0.3pp. Does your recommendation change?
- The PM says changing the metric now looks like you are protecting a feature. How do you answer?
- What would you need to see to agree the feature should be killed?
- 01
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
- 02
A supplier-growth team moved one market's inventory from 'strict' to 'flexible' cancellation_policy in fct_rate_availability_snapshot. Their dashboard shows bookings up 14% over eight weeks. Your read on fct_stay_night shows stayed nights flat, and traveller-initiated cancellation up from 18% to 31% in the same lead-time bucket. Their quarterly goal is booked nights, and the lead has already sent the 14% to their director. You have ten minutes in their weekly review. Prepare what you open with, what you concede, and what you will not soften.
- 03
A product manager wants to ship a filter-refinement panel that lets travellers re-sort and narrow results more easily. Early data shows it increases searches per trip_intent_key from 6.2 to 9.4. The team's headline metric is bookings divided by fct_search rows, which falls from 4.1% to 2.9% under the change. The PM reads this as the feature hurting conversion and wants to kill it. You think the denominator is wrong. Prepare how you make that case in a working session without winning on a technicality.
Is this an official Hopper interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Hopper. Rounds and questions reflect what candidates have reported, not a process Hopper has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Hopper Data Scientist interview process?
The process is rated as average to difficult, primarily due to the highly open-ended nature of the take-home challenge. Success requires a strong balance of technical execution, visual communication, and strategic product thinking.
PracHub interview research ↗What is the most common reason candidates fail the loop?
Most candidates struggle with the take-home challenge, either by failing to provide actionable business recommendations or by submitting poorly structured visualizations. Technical skills are necessary, but your business intuition and communication are what set you apart.
PracHub interview research ↗How much interaction will I have with executive leadership?
Quite a bit. The interview loop frequently includes rounds with the Chief Strategy Officer, VPs of Data Science, or Revenue Leaders. Hopper values data scientists who can hold their own in strategic business discussions.
PracHub interview research ↗Does Hopper provide feedback after the take-home challenge?
Historically, candidates have reported receiving limited detailed feedback upon rejection due to the high volume of applicants. It is highly recommended to self-review your work against professional standards before submitting.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22