A Data Scientist at Booking plays a pivotal role in leveraging data to drive decision-making and enhance user experience across the platform. This position is crucial as it directly impacts how products are tailored to meet customer needs and optimize operational efficiency. As a data scientist, you will be at the forefront of analysis, utilizing advanced statistical techniques and machine learning algorithms to extract insights from vast datasets. This influence extends to various teams and products, from improving search algorithms to personalizing customer recommendations, ultimately leading to increased customer satisfaction and loyalty.
Your work will involve complex problem-solving in a fast-paced environment, where the scale of operations presents unique challenges. You will engage with diverse data types and sources, contributing to the strategic direction of Booking by informing product development and marketing strategies. This role not only offers the opportunity to work with cutting-edge technologies but also places you in a collaborative environment where cross-functional teamwork is essential for success.
Candidates can expect to engage in high-impact projects, working alongside product managers, engineers, and marketing teams to enhance the overall customer journey. This is an exciting opportunity to make a significant difference in a leading global travel company.
Initial Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Interview
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Business Case Discussions
reportedA case has a fixed clock, and a good deal of what is being scored is how you spend it. Thirty to forty-five minutes buys one pass across the whole problem or a deep read of one part of it, and choosing between those is the work rather than a compromise forced on you. Announce the shape early: the structure you are using, the branch you think carries the decision, and what you are setting aside. An answer that is thorough for the first third and silent on the recommendation reads worse than one that is rougher throughout and lands.
What to demonstrate
- Whether a visible structure appears in the opening minutes and survives the rest of the case
- Whether the depth goes to the branch that carries the decision, rather than the branch you find most comfortable
- Whether you say what you are leaving out and why, instead of quietly omitting it and hoping nobody asks
How to prepare
- After each practice case, write down the branches you chose not to open and the reason for each, then check whether you said any of them out loud while the case was running. A branch you only cut privately reads to the interviewer as one you missed.
- Redo a case you have already worked in half the time, deciding in advance which single branch you keep, then compare which version a listener would find more useful.
- Write a two-sentence opening you can reuse, holding the restated question and your plan for the available time, and deliver it within the first ninety seconds of every practice run.
PracHub editorial advice for the preparation topics above.
Reading cancellation, completion or repeat rates on cohorts that have not matured
A cohort of bookings made last week for stays six months out cannot have cancelled at the check-in gate yet, so its cancellation rate is mechanically near zero and its completion rate mechanically near zero as well, in opposite directions. Comparing that cohort with a mature one is not a noisy comparison, it is a guaranteed wrong one, and the bias always makes the recent period look different in a way that invites a false story about a recent change. Because lead time is heavily right-skewed, the mean lead time is a bad maturity threshold; use the cohort's 95th percentile, or report a hazard at a fixed age (cancelled within k days of booking) with k capped at the youngest cohort's elapsed age. The same applies to repeat rate, where the honest answer is often that the cohort in question is not readable for another nine months.
Fitting demand or price elasticity on observed bookings, when availability and restrictions censor the data
Bookings equal the minimum of demand and what was actually sellable, so a sold-out date records the capacity, not the demand behind it, and the censoring is worst precisely on the highest-demand dates. Meanwhile price is set from a forecast of that same demand, so high-demand dates carry high prices and the raw correlation between price and bookings is biased toward zero and frequently comes out positive, which reads as 'raising price increases demand'. Restrictions compound it: a minimum-length-of-stay rule or a closed-to-arrival flag suppresses bookings with no price movement at all, so the effect lands on the price coefficient if the restriction is not in the model. Join fct_rate_availability_snapshot at snapshot_date equal to the search or booking date, restrict the estimation sample to unit-dates that were genuinely open, carry the restriction flags as controls, and lean on an instrument or a deliberate price experiment before quoting an elasticity.
Over-explaining the method and under-explaining the implication
Lead with the answer and what you would do about it, then give the approach when asked. Roughly one sentence of method per three of implication is the right ratio for a stakeholder-facing answer; the interviewer already knows what a regression is.
Naming a model class before naming the deployment constraints
Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Can you explain a machine learning algorithm you have implemented in t…
Can you explain a machine learning algorithm you have implemented in the past?
Approach
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Explain how you would implement a predictive model for customer churn.
Explain how you would implement a predictive model for customer churn.
Approach
- Say how the offline result would be validated online before it is trusted.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Expand bookings into stay nights without losing cents
You are given a pandas DataFrame bookings with columns booking_id, check_in_date, check_out_date, nights, unit_revenue_usd (the booking total, two decimal places), about 200,000 rows. Write a function returning stay_nights at one row per booking per stay date, with columns booking_id, stay_date, night_index (1-based) and unit_revenue_usd allocated across the nights so the per-booking sum equals the original to the cent. No per-row Python loop and no apply over rows. State your rule for the departure date.
Approach
- Confirm the grain rule first: a stay night is the arrival date through the night before check_out_date, so nights = (check_out_date - check_in_date).days and the departure date is never a row. Assert this against the stored
nightscolumn and report mismatches rather than trusting either side. - Build the expansion arithmetically instead of materialising date lists: counts = nights.to_numpy(); starts = np.repeat(np.cumsum(counts) - counts, counts); offsets = np.arange(counts.sum()) - starts. That gives a 0-based night offset per output row in one pass, and night_index = offsets + 1.
- Compute stay_date as np.repeat(check_in_date.to_numpy(), counts) + offsets.astype('timedelta64[D]'). This is O(total nights) with no grouping and no Python-level iteration.
- Allocate revenue in integer cents, not floats: cents = np.rint(unit_revenue_usd * 100); base = cents // counts; remainder = cents - base * counts; give the first
remaindernights of each booking one extra cent. Dividing floats by nights leaves sub-cent drift that re-sums to 99.99 or 100.01 on a large fraction of bookings. - Verify by grouping the output back to booking_id and comparing integer cents to the input, then state which downstream numbers depend on this reconciliation (net revenue per sellable night, take rate, contribution margin per stayed night all read the night grain).
Worked solution 25 min
- Assert (check_out_date - check_in_date).dt.days equals nights; raise on any mismatch and print the offending booking_ids rather than silently trusting the column.
- Build counts, starts and offsets with np.repeat and np.arange as above; total output rows equals counts.sum().
- Derive stay_date by adding the day offsets to the repeated check_in_date, and night_index as offsets + 1.
- Allocate integer cents with floor division plus largest-remainder top-up on the first
remaindernights of each booking. - Reconcile: group the output by booking_id, sum the cents, compare exactly to the input cents, assert zero differing rows.
Follow-up
- The booking is later amended from three nights to five. What has to happen to the already-written night rows, and which of them may not be rewritten?
- Nights of one booking fall in two calendar months and two tax jurisdictions. Does your equal-split allocation still defend the monthly revenue number, and when would you weight nights by the nightly listed price instead?
Compare booking pace against last year aligned on days out
fct_rate_availability_snapshot carries units_sold_to_date, the cumulative pickup for a unit-date as at snapshot_date, and lead_time_days, which is stay_date minus snapshot_date. For one destination_market_id and one target stay week, build the booking curve: for each lead_time_days from 120 down to 0, the total units sold across the market's units and the seven-day pickup. Join the equivalent stay week last year aligned on lead_time_days rather than on calendar date, and return both curves side by side.
Approach
- Restrict to the market's supply units and the seven stay_dates of the target week, then aggregate to lead_time_days by summing units_sold_to_date across units and across those seven stay_dates. This is a cross-section at a fixed days-out, so the snapshot_date differs across the week's stay dates, which is exactly what aligning on days out means.
- Repeat for last year's aligned stay week, chosen on the market calendar. A 364-day offset preserves day of week but drifts a moving holiday; a same-date offset preserves the date and breaks day of week. State which you chose, and use market_holiday_flag on fct_stay_night to check whether a holiday sits inside one week and not the other.
- Compute seven-day pickup as value minus LAG(value, 7) OVER (ORDER BY lead_time_days DESC). Ordering descending makes time run forward, because days out counts down towards arrival. Every value is a rolling window and adjacent rows share six of their seven days, so this column is not additive and must never be summed to recover total pickup.
- Join the two curves on lead_time_days with a FULL OUTER JOIN so a days-out point missing on one side stays visible as a NULL rather than dropping the row and silently shortening the comparison.
- Index each curve to its own final on-hand value when the two weeks differ in sellable size, so the comparison reads as pace rather than level.
Worked solution 45 min
- Pull the raw snapshot rows for one supply unit and one stay_date ordered by lead_time_days, and confirm units_sold_to_date is already cumulative before writing any SUM over time.
- Aggregate to the market-week cross-section and check monotonicity as days out falls.
- Add the LAG-based seven-day pickup and validate a single row straight against the cumulative column: pickup at lead_time_days 30 must equal cumulative at 30 minus cumulative at 37. Do not validate it by summing the column, which the overlap makes meaningless.
- Build last year's curve, align on lead_time_days, and note where the holiday falls in each week before drawing any conclusion.
- Join, index each series to its final value, and report the absolute and the indexed gap together.
Follow-up
- At 30 days out you are 8% behind last year. What else do you need before telling revenue management to cut price?
- Two new units joined this market in March. What does that do to this comparison, and how would you control for it?
- The curve at days out zero sits above the stayed-night count for that week. Is that a bug, and how would you show which it is?
Intent-to-booking conversion within a seven-day attribution window
fct_search carries trip_intent_key, searched_at_utc, device_type, lead_time_days, results_returned_count and is_bot_flagged; fct_booking carries booking_id, trip_intent_key, booked_at_utc and booking_status. For intents whose first qualifying search falls in one month, compute the intent-to-booking rate: an intent converts if any booking shares its trip_intent_key and was booked within seven days of that first search, counting bookings that later cancel. Qualifying searches are is_bot_flagged = FALSE with results_returned_count > 0. Report numerator, denominator and rate by device_type and lead-time bucket.
Approach
- Filter to qualifying searches first, then collapse to one row per trip_intent_key using MIN(searched_at_utc), and state the convention plainly: an intent whose earliest search returned nothing is cohorted on its first search that did return results.
- Carry device_type and lead_time_days from that first qualifying search, picked with ROW_NUMBER or FIRST_VALUE. One intent routinely spans mobile and desktop, and attributing it to the device of the converting search inflates whichever device closes.
- Aggregate bookings to one row per trip_intent_key with MIN(booked_at_utc) before joining. Joining raw bookings makes an intent with three bookings count three times in the numerator and, once grouped, in the denominator too.
- LEFT JOIN the booking aggregate onto the intent set and flag conversion where booked_at_utc falls within [first_search_at, first_search_at + 7 days]. Keep cancelled and no-show bookings in: this metric measures demand conversion, not delivery.
- Group by device and bucket, summing the flag and counting intents, and compute the rate as the ratio of those two sums rather than as an average of per-cell rates.
Follow-up
- Mobile converts at half the desktop rate. How much of that gap is mix, and what decomposition would you run to show it?
- How do you hold the seven-day window open without restating last week's published number every morning?
- A ranking change cuts searches per intent by 20% and leaves bookings flat. What happens to this metric, and what happens to a search-denominated version of it?
Given a dataset of customer transactions, how would you identify patte…
Given a dataset of customer transactions, how would you identify patterns in purchasing behavior?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Explain how you would approach a project that requires cross-team coll…
Explain how you would approach a project that requires cross-team collaboration.
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How do you prioritize your workload when dealing with multiple project…
How do you prioritize your workload when dealing with multiple projects?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What methods would you use to assess the success of a marketing campai…
What methods would you use to assess the success of a marketing campaign?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you design an experiment to test a new feature on the Bookin…
How would you design an experiment to test a new feature on the Booking platform?
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
What statistical methods do you commonly use for A/B testing, and why?
What statistical methods do you commonly use for A/B testing, and why?
Approach
- Say whether units interfere with each other, and switch design if they do.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
How do you approach data visualization, and what tools do you use?
How do you approach data visualization, and what tools do you use?
Approach
- Say what you would check first and why it is the highest-information step.
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Estimate a market-level policy rollout without randomisation
A mandatory change to displayed cancellation terms takes effect in six destination markets on dates set outside the business; it cannot be randomised or withheld. You have 104 weeks of market-by-week history for 186 markets on completed stay-nights and net revenue per sellable night. Four markets adopt on one date and two adopt three months later. Specify the estimator, how you test the identifying assumption, how you obtain standard errors from six treated units, and which date axis the outcome sits on.
Approach
- Settle the date axis first. The policy acts on the booking date, but completed stay-nights are consumed later, so the stay-axis effect arrives smeared across the lead-time distribution. Run a booking-axis outcome for the sharp read and a stay-axis outcome for the true one, and treat a knife-edge break on the stay axis as evidence of a data problem rather than of the policy.
- Reject plain two-way fixed effects because adoption is staggered: already-treated markets serve as controls for the later adopters and receive negative weights, which can reverse the sign when effects are dynamic. Use a group-time ATT estimator (Callaway-Sant'Anna or Sun-Abraham) with never-treated markets as the comparison group, aggregated into an event-time path.
- Test parallel trends with an event study on leads from week -12 to week -1 and run the joint test on those coefficients rather than inspecting the plot. Add a sensitivity analysis that reports how large a post-treatment violation of parallel trends would have to be to overturn the conclusion, since passing a pre-trend test is weak evidence.
- Fix the inference for six treated clusters. Cluster-robust standard errors are badly anti-conservative well above six clusters, so use randomisation inference: reassign the treatment labels with their adoption dates across the 186 markets many times and place the observed estimate in that null distribution, reporting the achievable p-value granularity alongside it.
- Cross-check with synthetic control. Build a donor-weighted synthetic counterpart for each treated market, require a small pre-period RMSPE, and evaluate significance with the post-to-pre RMSPE ratio against the placebo distribution over donors. With six treated units, synthetic difference-in-differences is the natural pooled version.
- Verify the mechanism rather than only the headline: the effect should show up in rate_plan mix and in the traveller-initiated cancellation rate by lead-time bucket, and an effect that appears only in the aggregate is a sign of a confounded comparison.
Worked solution 45 min
- Build the market-by-week panel with both outcomes and both date axes, flagging each market's adoption week and the never-treated set.
- Estimate group-time ATTs for the two adoption cohorts separately, then aggregate into an event-time path from week -12 to week +12.
- Run the joint pre-trend test on the lead coefficients and record the p-value as a stated assumption check, not as proof.
- Run randomisation inference by permuting the six treatment assignments and their dates across the 186 markets, and report the observed estimate's rank in that distribution.
- Fit the synthetic control per treated market, report pre-period RMSPE and the post-to-pre ratio against donor placebos, and compare the pooled synthetic DiD estimate with the group-time aggregate.
Follow-up
- One of the six markets is far larger than the others. What does that do to the estimator and to the randomisation-inference null?
- Suppose all six had adopted on the same date. What design would you use instead, and what do you lose?
- The ops team offers to delay rollout in two more markets by a month. What exactly would you ask them to delay, and why is that request worth making?
Blended daily rate fell while every segment's rate rose
Blended ADR for last month, computed from fct_stay_night as SUM(unit_revenue_usd) / COUNT(*) over night_status = 'stayed' at a fixed reference FX rate, fell 4.2% year over year. Cut by destination_market_id, by property_class and by rate_plan, every segment's ADR rose. Commercial leadership is drafting a note about discounting. Using fct_stay_night joined to dim_supply_unit (property_class, destination_market_id) and fct_booking (rate_plan, lead_time_days), quantify how much of the -4.2% is mix and how much is within-segment, and show that your segmentation choice is not doing the work.
Approach
- Write the identity before computing anything: blended ADR equals the sum over segments of w_s times ADR_s, where w_s is segment s's share of stayed nights. The change then decomposes into a mix term, sum of (w_s1 - w_s0) times ADR_s0; a within-segment term, sum of w_s0 times (ADR_s1 - ADR_s0); and an interaction term. Report the interaction rather than folding it silently into one side.
- Choose the segmentation from the mechanism rather than from convenience: market by property_class by rate_plan by lead-time bucket. A coarse cut leaves mix hiding inside segments, which makes the decomposition understate exactly the term you are trying to size.
- Compute w_s on stayed nights, the same denominator the metric uses. Weighting by bookings or by revenue produces terms that do not re-aggregate to the published blended ADR, and the reconciliation failure is how you will find out.
- Test that the answer is not an artefact of the cut. Rerun at two or three segmentation depths and report the spread in the mix share; if it swings wildly, the finest cut is fitting thin cells and the cell counts belong next to the numbers.
- Only now ask what moved the weights, and look for a cause with a date attached: a market that grew supply, a channel that brought shorter lead times, a seasonal shift in party size or length of stay.
Follow-up
- You have the decomposition. What do you actually recommend, given within-segment rate is up?
- The mix moved because one low-ADR market grew 40% year over year. Is that a problem to fix or the plan working as intended?
- How do you keep this decomposition honest when a segment exists this year that did not exist last year, so w_s0 is zero for it?
For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Build a fixture you can check answers against
- Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
- Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
- Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.
Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Joins, filters and NULL semantics
- Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
- Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
- Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.
Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Window functions and frames
- Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
- Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
- Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.
Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The four analytical query patterns
- Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
- Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
- Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.
Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Write SQL the way you will have to write it live
- Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
- Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
- Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.
Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.
Practice prompt ↗Practice prompt ↗06One day for everything that is not SQL
- Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
- Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
- Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.
Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.
Practice prompt ↗Practice prompt ↗07Full loop rehearsal
- Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
- Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
- Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.
Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
How do you ensure that your work aligns with the company's goals and m…
How do you ensure that your work aligns with the company's goals and mission?
Approach
- Close with what you would do differently, concretely.
- State the situation in two sentences and spend the rest on your reasoning.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Retract an occupancy comparison after the decision shipped
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
Approach
- Quantify the error before telling anyone, because the first question will be 'how wrong'. Recompute both markets on sellable nights, which excludes units_blocked, and report the corrected gap and its sign, not just that the original was wrong.
- Separate the numerical error from the decision error. The comparison was invalid, but the spend move may still have been right; establishing whether the decision would have flipped is a different analysis and it is the one the business needs.
- Tell the decision-maker directly and first, before it appears in a dashboard or a peer surfaces it. Order matters because being told by a third party converts a mistake into a credibility problem.
- Write the note with the correction, the size, the decision implication, and the reversal cost in that order. Three weeks of moved spend has a real cost to undo, and a correction that does not price the reversal forces the reader to do the work you skipped.
- Name the specific control that would have caught it and put it in place in the same note: an assertion in the query that both arms use the same denominator expression, and the denominator named in the chart title. A retraction without a mechanism reads as an apology rather than a fix.
Follow-up
- The corrected numbers still support the original decision. Do you still send the note, and does it read differently?
- Your manager suggests quietly fixing the dashboard and not raising it. How do you respond?
- What would you have had to do differently three weeks ago, in the query itself, rather than in your review habits?
Give an executive one number and its honest interval
A director must commit contact-centre staffing for a peak stay-week 60 days out. Your pickup forecast from fct_rate_availability_snapshot gives 42,000 stayed nights for that week, with an 80% interval of 36,500 to 47,000. Back-testing shows the interval only narrows materially inside 30 days out. The director has said twice that they want one number and do not want to hear about confidence intervals. You have a five-minute slot and one slide. Prepare what goes on it and what you say.
Approach
- Work out the decision before the number: staffing is asymmetric, because under-staffing a peak week costs servicing failures and cancellations while over-staffing costs idle hours. Find out which side is more expensive, because the whole answer is which end of the interval to plan against.
- Convert the interval into two staffing levels rather than two stay-night counts. An executive cannot act on 36,500 to 47,000; they can act on 'staff for 44,000 now, with a named trigger to add or release capacity'.
- Give the one number they asked for, and attach the trigger to it rather than a caveat: plan at the level implied by the more expensive error, then re-read pickup at 30 days out, which back-testing says is where the interval actually moves.
- Say what the uncertainty is made of in one line each, using the domain's own vocabulary: pickup still to come, lead-time mix for that market, and whether the week straddles a moving holiday that shifts against last year.
- Close with the decision the interval does not affect, so the director knows the width is not an excuse: if every point in the interval implies the same staffing tier, say so and stop talking about the interval.
Follow-up
- The week lands 6,000 nights below your point forecast. How do you handle the next forecast conversation with this director?
- They ask you to just give them the midpoint and skip the trigger. What do you do?
- How would you decide whether an 80% interval is the right one to show rather than a 50% or a 95%?
- 01
How do you ensure that your work aligns with the company's goals and mission?
- 02
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
- 03
A director must commit contact-centre staffing for a peak stay-week 60 days out. Your pickup forecast from fct_rate_availability_snapshot gives 42,000 stayed nights for that week, with an 80% interval of 36,500 to 47,000. Back-testing shows the interval only narrows materially inside 30 days out. The director has said twice that they want one number and do not want to hear about confidence intervals. You have a five-minute slot and one slide. Prepare what goes on it and what you say.
Is this an official Booking interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Booking. Rounds and questions reflect what candidates have reported, not a process Booking has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the typical difficulty level of the interviews?
The interviews for the Data Scientist role at Booking can be challenging, especially the technical portions. Candidates should expect to demonstrate both their analytical skills and their cultural fit. Preparing thoroughly can significantly enhance your chances of success.
PracHub interview research ↗How long does the interview process usually take?
The timeline from initial screening to offer can vary, but candidates often report a process that spans several weeks. It's important to stay proactive in your communication with recruiters during this time.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates often demonstrate a mix of strong technical skills, effective communication, and a clear alignment with Booking's values. Being able to articulate how your experience relates to the company's mission is key.
PracHub interview research ↗How does Booking support professional development?
Booking values continual learning and development. Employees often have access to training resources and opportunities to attend industry conferences, which help them stay updated with the latest trends and technologies in data science.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22