A Data Scientist at Priceline serves as a vital bridge between complex data architecture and high-stakes business strategy. In the highly competitive travel and hospitality industry, your work directly influences the algorithms that power travel recommendations, pricing strategies, and user conversion funnels. You are not just building models; you are solving real-world problems that dictate how millions of users discover and book their next trip.
Success in this role requires a blend of rigorous technical execution and a sharp product mindset. You will be expected to translate ambiguous business questions into measurable analytical frameworks, leveraging data to drive decision-making across product, marketing, and engineering teams. Whether you are optimizing a recommendation engine or designing a high-impact A/B test, your contributions are expected to be both scientifically sound and commercially impactful.
Recruiter Screening
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Take-Home Assignment
reportedYour submission is read asynchronously by someone who cannot ask you a clarifying question, and who will usually skim it before reading it properly. That changes what good looks like. The conclusion belongs near the top with the supporting analysis beneath it, and the code should run start to finish on a clean machine without a manual step you forgot to document. A reviewer forced to reconstruct your reasoning from the order of notebook cells is already discounting the work. The submissions that land are the ones where a busy reader gets the answer immediately and can verify it if they want to.
What to demonstrate
- Whether the answer arrives before the methodology, so a reader who stops after the first page still has the recommendation
- Whether the code runs end to end from the submitted files, with dependencies and data paths declared rather than assumed
- Whether each chart is legible on its own, carrying axis labels and units, and exists to support a claim made in the text
How to prepare
- Restructure a past analysis so the opening paragraph holds the recommendation and the number behind it, then check that nothing later in the document quietly contradicts it
- Copy your own submission into an empty directory, run it in a clean environment, and fix everything that breaks; hidden local state is caught here or by the reviewer
- For every chart, write the one sentence it is meant to prove, and delete the chart if you cannot write that sentence
Final Round Interviews
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub editorial advice for the preparation topics above.
Fitting demand or price elasticity on observed bookings, when availability and restrictions censor the data
Bookings equal the minimum of demand and what was actually sellable, so a sold-out date records the capacity, not the demand behind it, and the censoring is worst precisely on the highest-demand dates. Meanwhile price is set from a forecast of that same demand, so high-demand dates carry high prices and the raw correlation between price and bookings is biased toward zero and frequently comes out positive, which reads as 'raising price increases demand'. Restrictions compound it: a minimum-length-of-stay rule or a closed-to-arrival flag suppresses bookings with no price movement at all, so the effect lands on the price coefficient if the restriction is not in the model. Join fct_rate_availability_snapshot at snapshot_date equal to the search or booking date, restrict the estimation sample to unit-dates that were genuinely open, carry the restriction flags as controls, and lean on an instrument or a deliberate price experiment before quoting an elasticity.
Reporting on the booking date when the question is about the stay date, or the reverse
The same booking contributes to demand in one period and to revenue, occupancy and supplier payout in another, so almost every headline number has two defensible values that can differ by tens of percent. A promotion that moves bookings this week for stays in March produces a revenue 'drop' on the stay axis and a spike on the booking axis, and both are real. Cancellations make it worse, because a cancellation dated today removes revenue from a stay date months out and will silently restate a period that was already reported closed. The only defence is to name the date axis in the title of every chart and every metric definition, and to keep a pace view (bookings-on-hand for a future stay date, by days out) as a separate artefact rather than trying to serve both from one table.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
Stopping an experiment the moment it crosses significance
Fix the sample size or duration before launch, or use a method built for continuous monitoring such as a sequential test, always-valid confidence intervals, or group-sequential boundaries. Repeatedly checking a fixed-horizon p-value against 0.05 pushes the real false-positive rate well above 5 percent.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Expand bookings into stay nights without losing cents
You are given a pandas DataFrame bookings with columns booking_id, check_in_date, check_out_date, nights, unit_revenue_usd (the booking total, two decimal places), about 200,000 rows. Write a function returning stay_nights at one row per booking per stay date, with columns booking_id, stay_date, night_index (1-based) and unit_revenue_usd allocated across the nights so the per-booking sum equals the original to the cent. No per-row Python loop and no apply over rows. State your rule for the departure date.
Approach
- Confirm the grain rule first: a stay night is the arrival date through the night before check_out_date, so nights = (check_out_date - check_in_date).days and the departure date is never a row. Assert this against the stored
nightscolumn and report mismatches rather than trusting either side. - Build the expansion arithmetically instead of materialising date lists: counts = nights.to_numpy(); starts = np.repeat(np.cumsum(counts) - counts, counts); offsets = np.arange(counts.sum()) - starts. That gives a 0-based night offset per output row in one pass, and night_index = offsets + 1.
- Compute stay_date as np.repeat(check_in_date.to_numpy(), counts) + offsets.astype('timedelta64[D]'). This is O(total nights) with no grouping and no Python-level iteration.
- Allocate revenue in integer cents, not floats: cents = np.rint(unit_revenue_usd * 100); base = cents // counts; remainder = cents - base * counts; give the first
remaindernights of each booking one extra cent. Dividing floats by nights leaves sub-cent drift that re-sums to 99.99 or 100.01 on a large fraction of bookings. - Verify by grouping the output back to booking_id and comparing integer cents to the input, then state which downstream numbers depend on this reconciliation (net revenue per sellable night, take rate, contribution margin per stayed night all read the night grain).
Follow-up
- The booking is later amended from three nights to five. What has to happen to the already-written night rows, and which of them may not be rewritten?
- Nights of one booking fall in two calendar months and two tax jurisdictions. Does your equal-split allocation still defend the monthly revenue number, and when would you weight nights by the nightly listed price instead?
Simulate an overbooking limit against rate-plan no-show risk
One property, one stay date, 40 identical units. Historical traveller no-show plus same-day cancellation rates by rate plan: non-refundable 2%, partially refundable 6%, refundable 11%. Confirmed bookings arrive in the mix 50% non-refundable, 20% partially, 30% refundable. Relocating a displaced guest costs 220 USD; an unsold unit costs 180 USD of foregone rate. Simulate to choose the number of bookings N (N at least 40) that minimises expected cost, report the Monte Carlo standard error at the chosen N, and state the assumption your simulation makes about independence.
Approach
- Write the cost function explicitly before simulating: cost = 220 * max(arrivals - 40, 0) + 180 * max(40 - arrivals, 0). The asymmetry is the whole problem, so name the ratio 220/180 early; it is what pushes the optimum past the point where expected arrivals equal capacity.
- Draw a single uniform matrix of shape (n_sims, N_max) once and assign rate plans to booking slots in an interleaved order, so that any prefix of N columns preserves the 50/20/30 mix. Reusing those same draws for every candidate N is common random numbers: it makes the cost differences between adjacent N far more precise than the level of cost at any one N.
- Sweep N from 40 to 48, compute mean cost and the standard error as std/sqrt(n_sims), and only then declare an argmin. If adjacent N values differ by less than a couple of standard errors, say so instead of picking one.
- Sanity-check the simulation against arithmetic: expected shows per booking is 0.500.98 + 0.200.94 + 0.30*0.89 = 0.945, and a normal approximation to the Poisson-binomial with that mean and variance should land on the same optimum.
- State the independence assumption and where it fails: a weather event, a flight disruption or a group booking split across several reservations correlates no-shows, which fattens both tails and moves the optimum toward capacity. Independent Bernoullis understate the chance of a costly relocation day.
Follow-up
- How does the optimum move if relocation cost rises to 600 USD, and what is the general rule relating the two costs to the optimal exceedance probability?
- The no-show rates are estimated from 800 historical nights. How would you propagate that estimation uncertainty into the choice of N?
- Give a design that would let you measure the true relocation cost, including the parts that do not appear on an invoice.
Build a censoring-aware pickup curve from availability snapshots
snap has supply_unit_id, stay_date, snapshot_date, units_sold_to_date, units_available, is_closed_to_arrival, min_length_of_stay, lead_time_days. Snapshots are irregular: a unit-date is not captured every day. Produce, per destination market, the mean cumulative share of final pickup on hand at each whole days-out value from 90 down to 0, computed only over unit-dates that were genuinely open at that horizon, plus a companion column giving the share of unit-dates excluded at each horizon. No per-unit Python loop. Say which direction the exclusion biases the curve.
Approach
- Define final pickup as units_sold_to_date at the smallest observed lead_time_days for that (unit, stay_date), and flag the caveat: it still contains bookings that later cancel, so this is a pace curve, not a stayed-nights curve, and it must not be relabelled as the latter. units_sold_to_date is cumulative net of cancellations, so it can fall between two snapshots; an earlier snapshot can then sit above the final-pickup denominator and give a share above 1.0 at a positive horizon with the grid pick running correctly. Do not clip that to 1.0 — count it and report the affected unit-date share, because clipping hides a real cancellation and a wrong fill direction behind the same flat ceiling.
- Put the irregular snapshots onto the regular 0..90 days-out grid with merge_asof on lead_time_days, both sides sorted ascending, direction='forward' and a
byof (supply_unit_id, stay_date). Forward on an ascending lead-time key selects the nearest snapshot with lead_time at least the grid value, which is the most recent capture that was still that far out. - Be explicit about the fill direction, because this is where the leak lives. On a lead-time-ascending index, a plain forward fill propagates small lead times into large ones, pushing near-final pickup back into the 60-day horizon and flattening every curve toward 100%. The correct fill on that index is backward.
- Mark a unit-date censored at a horizon when units_available is 0 or is_closed_to_arrival is true or min_length_of_stay exceeds 1, and compute the mean share over the uncensored rows only. Count and report the excluded share per horizon in the same frame so the reader sees the sample changing under the curve.
- State the bias direction: censoring is correlated with demand, so the excluded unit-dates are disproportionately the fast-selling weekend and holiday dates. Dropping them biases the surviving curve toward slow-selling inventory, and including them biases it the other way because a sold-out date sits at 100% of final pickup early. Neither version is the demand curve, and that is the point worth saying out loud.
Worked solution 45 min
- Compute final pickup per (supply_unit_id, stay_date) as units_sold_to_date at the minimum lead_time_days; drop unit-dates whose smallest observed lead_time exceeds 3 days, since their final value is unknown.
- Build the cross of unit-dates with the 0..90 grid, then merge_asof on lead_time_days with by=['supply_unit_id','stay_date'] and direction='forward'.
- Derive share = units_sold_to_date / final_pickup, guarding final_pickup == 0 by excluding those unit-dates and reporting how many.
- Build the censored mask from units_available == 0, is_closed_to_arrival, and min_length_of_stay > 1; aggregate mean share over uncensored rows and the excluded fraction over all rows, both grouped by market and days_out.
- Plot or tabulate both series together and write the one-sentence bias statement.
Follow-up
- How would you use this curve as an input to a stay-week forecast, and at which days-out horizon does it stop carrying information?
- Two markets have visibly different curves. What does that imply for how far out each one's price should be set, and for the confidence interval on their forecasts?
- Sketch the survival-analysis version of this problem and say what it buys you over the exclusion approach.
How would you structure a query to identify top-performing travel dest…
How would you structure a query to identify top-performing travel destinations by user segment?
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Explain the difference between `RANK`, `DENSE_RANK`, and `ROW_NUMBER`.
Explain the difference between RANK, DENSE_RANK, and ROW_NUMBER.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
How do you utilize SQL window functions to calculate running totals or…
How do you utilize SQL window functions to calculate running totals or period-over-period growth?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Pick the last availability snapshot before each stay date
fct_rate_availability_snapshot is keyed on (supply_unit_id, stay_date, snapshot_date) and carries units_available, units_sold_to_date, units_blocked, listed_price_usd, is_closed_to_arrival and min_length_of_stay. For every supply_unit_id and every stay_date in one target month, return exactly one row: the values from the latest snapshot_date that is on or before that stay_date. Output supply_unit_id, stay_date, snapshot_date, units_available, units_sold_to_date and listed_price_usd. A single global maximum snapshot_date is not acceptable.
Approach
- Restrict to the target stay_date range and to snapshot_date <= stay_date inside the same CTE; that predicate is evaluated per row, which is what makes this an as-of pick rather than a latest-snapshot pick.
- Rank with ROW_NUMBER() OVER (PARTITION BY supply_unit_id, stay_date ORDER BY snapshot_date DESC) and keep rank 1. RANK() would return two rows if the table ever holds a duplicate snapshot_date; ROW_NUMBER guarantees one.
- Filter on the rank in an outer query, or use DISTINCT ON / QUALIFY where the engine supports it. A window function cannot be referenced in the WHERE clause of the SELECT that computes it.
- Leave unit-dates with no qualifying snapshot out of the result and say so explicitly; if the deliverable needs them present, LEFT JOIN from a unit-by-date spine so the absence surfaces as a NULL row instead of a missing one.
- Keep snapshot_date in the output, because every downstream consumer needs to know how stale the as-of value is before acting on it.
Worked solution 20 min
- Count distinct (supply_unit_id, stay_date) pairs in the window; that is the row count the final query must return, less any unit-dates with no snapshot on or before arrival.
- Write the ranked CTE, then select rank 1 from an outer query.
- Assert across the result that no snapshot_date exceeds its stay_date.
- Spot-check one unit-date against its raw snapshot rows to confirm the chosen row really is the last capture before arrival.
Follow-up
- Now pick the snapshot as at exactly 30 days before arrival instead. What changes in the query, and which of the two do you use for a pace report?
- The table is partitioned on snapshot_date. How do you keep partition pruning while still ranking within each unit-date?
- You find two rows sharing a snapshot_date for one unit-date. What does that mean about the capture pipeline, and how do you break the tie deterministically?
How would you define the success metrics for a new hotel search featur…
How would you define the success metrics for a new hotel search feature?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Describe a project where you had to pivot your approach due to shiftin…
Describe a project where you had to pivot your approach due to shifting business priorities.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
If a key conversion metric drops suddenly, what is your systematic app…
If a key conversion metric drops suddenly, what is your systematic approach to diagnosing the root cause?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
How do you determine the required sample size for statistical signific…
How do you determine the required sample size for statistical significance?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Decide the analysis before seeing data, including how long it runs and when you look.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
What are the most common experimentation pitfalls that lead to false p…
What are the most common experimentation pitfalls that lead to false positives?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
How would you design an Amazon-like recommendation system for travel b…
How would you design an Amazon-like recommendation system for travel bookings?
Approach
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Design metrics for a campaign that opens supplier availability
A supplier team runs a campaign asking individual_host and independent supply to open blocked dates for sale. Success is currently measured as sellable nights added, computed from fct_rate_availability_snapshot as units_available plus units_sold_to_date. Net revenue per sellable night is a company guardrail. Opening dates enlarges that guardrail's denominator immediately and its numerator only if the nights sell. Define the primary metric, resolve the mechanical conflict with the guardrail, and name how a supplier can move the campaign metric with no change in real supply.
Approach
- Name the mechanical conflict exactly: net revenue per sellable night equals ADR times occupancy over the same set, so opening nights that do not sell lowers it by construction. The guardrail will go red on a campaign that is working, and holding it flat would mean opening only dates that were already going to sell.
- Resolve it by scoping rather than by weakening: hold net revenue per sellable night on the pre-campaign cohort of unit-dates, those open before launch, and report newly opened unit-dates as a separate cohort with their own occupancy and ADR, so dilution is not mistaken for damage and cannibalisation into existing dates stays visible.
- Set the primary at the outcome rather than the input: incremental stayed nights and incremental platform_net_revenue_usd from newly opened unit-dates, differenced against a matched control set of units that did not open dates, because sellable nights added is an input the supplier controls directly.
- Name the gaming concretely: a supplier can open dates at a price that will never clear, or open and re-block after the snapshot that scores the campaign. Score only unit-dates that stayed open continuously through to their stay date, and require listed_price_usd within a band of that unit's own trailing median.
- Add the guardrails that catch the real failure modes: 90-day re-blocking rate, supplier retention, supplier-caused cancellation rate on newly opened dates, and occupancy on the pre-campaign cohort to detect demand simply moving off dates that were already selling.
Worked solution 45 min
- Freeze two cohorts of (supply_unit_id, stay_date) pairs at the snapshot_date immediately before launch: dates already open, and dates opened during the campaign.
- Compute occupancy, ADR and net revenue per sellable night separately within each cohort, and confirm the identity revenue per sellable night = ADR times occupancy holds within each before combining anything.
- Build a control group of units that opened no dates, matched on destination_market_id, property_class, supplier_type and pre-period occupancy, and difference the newly opened cohort against it.
- Run the re-blocking check: for each opened unit-date, verify units_blocked stayed at zero through the stay date, count the reversions, and exclude them from the success numerator.
- Report pre-campaign cohort occupancy as the cannibalisation read, and state how much of the incremental stayed nights survives if the entire decline there is attributed to the campaign.
Follow-up
- Newly opened nights sell at 40 percent occupancy against 72 percent on existing dates. Is the campaign working?
- How would you randomise this to separate cannibalisation from incremental nights, given that units in one market compete for the same travellers?
- Which single number would you put in front of finance, and on which date axis?
Blended daily rate fell while every segment's rate rose
Blended ADR for last month, computed from fct_stay_night as SUM(unit_revenue_usd) / COUNT(*) over night_status = 'stayed' at a fixed reference FX rate, fell 4.2% year over year. Cut by destination_market_id, by property_class and by rate_plan, every segment's ADR rose. Commercial leadership is drafting a note about discounting. Using fct_stay_night joined to dim_supply_unit (property_class, destination_market_id) and fct_booking (rate_plan, lead_time_days), quantify how much of the -4.2% is mix and how much is within-segment, and show that your segmentation choice is not doing the work.
Approach
- Write the identity before computing anything: blended ADR equals the sum over segments of w_s times ADR_s, where w_s is segment s's share of stayed nights. The change then decomposes into a mix term, sum of (w_s1 - w_s0) times ADR_s0; a within-segment term, sum of w_s0 times (ADR_s1 - ADR_s0); and an interaction term. Report the interaction rather than folding it silently into one side.
- Choose the segmentation from the mechanism rather than from convenience: market by property_class by rate_plan by lead-time bucket. A coarse cut leaves mix hiding inside segments, which makes the decomposition understate exactly the term you are trying to size.
- Compute w_s on stayed nights, the same denominator the metric uses. Weighting by bookings or by revenue produces terms that do not re-aggregate to the published blended ADR, and the reconciliation failure is how you will find out.
- Test that the answer is not an artefact of the cut. Rerun at two or three segmentation depths and report the spread in the mix share; if it swings wildly, the finest cut is fitting thin cells and the cell counts belong next to the numbers.
- Only now ask what moved the weights, and look for a cause with a date attached: a market that grew supply, a channel that brought shorter lead times, a seasonal shift in party size or length of stay.
Follow-up
- You have the decomposition. What do you actually recommend, given within-segment rate is up?
- The mix moved because one low-ADR market grew 40% year over year. Is that a problem to fix or the plan working as intended?
- How do you keep this decomposition honest when a segment exists this year that did not exist last year, so w_s0 is zero for it?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Saying no well is a senior skill and it is rarely rehearsed. Think of a time you told someone their analysis was not worth doing, or that the experiment could not answer their question at the sample size available. Explain what you offered instead. Refusal without an alternative reads as obstruction rather than judgement.
Tell me about a time you disagreed with a product manager’s decision; …
Tell me about a time you disagreed with a product manager’s decision; how did you handle it?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What did you decide not to do, and why?
- How did you know the outcome was caused by your change?
Defend a cancellation finding against the team it damages
A supplier-growth team moved one market's inventory from 'strict' to 'flexible' cancellation_policy in fct_rate_availability_snapshot. Their dashboard shows bookings up 14% over eight weeks. Your read on fct_stay_night shows stayed nights flat, and traveller-initiated cancellation up from 18% to 31% in the same lead-time bucket. Their quarterly goal is booked nights, and the lead has already sent the 14% to their director. You have ten minutes in their weekly review. Prepare what you open with, what you concede, and what you will not soften.
Approach
- Before the meeting, rebuild both periods as a hazard at a fixed age: share cancelled within k days of booked_at_utc, with k capped at the elapsed age of the youngest cohort. The flexible-policy cohort is younger, so a raw cancellation rate would be low for maturity reasons alone, and presenting that comparison hands the room a correct objection that kills the finding on its first sentence.
- Open by conceding the part that is true and theirs: bookings did rise 14%, the campaign did what it was designed to do on the booking axis. Naming their win first removes the reading that you are attacking the team rather than the metric.
- State the disagreement as an axis disagreement, not a competence one: booked nights is counted on booked_at_utc, stayed nights on stay_date, and the gap between them is exactly what a cancellation-policy change moves. Show the two series on one chart with both axes labelled.
- Quantify the delivered outcome in their own units so the tradeoff is arithmetic rather than opinion: stayed nights flat means the incremental bookings cancelled at roughly the rate that absorbs the whole 14%, and contribution margin per stayed night absorbed the servicing and payment-processing cost of the bookings that did not convert.
- Offer a route that keeps their goal intact: propose booked-nights-net-of-cancellation as the team's tracked number, or a stay-date readout held until the cohort matures, and say which you would commit to defending upward on their behalf.
Follow-up
- The lead says your cancellation cohort is not mature enough to compare. Walk me through the exact calculation that makes it comparable.
- Their director asks you directly whether the campaign should be rolled back. What do you say, and what would change your answer?
- How would you have set this up eight weeks ago so this conversation never happened?
Refuse the denominator that flatters a launch
A launch review is tomorrow. Occupancy computed on sellable nights, SUM(units_sold_to_date + units_available) from fct_rate_availability_snapshot, shows the launch market up 0.4pp. Computed on physical capacity, capacity_units from dim_supply_unit, it shows up 2.1pp, because individual_host suppliers blocked dates during the period and units_blocked is excluded from the first denominator but inside the second. The launch owner asks you to use the second, noting it is a documented convention. Prepare your response and what you put in the review document.
Approach
- Confirm the mechanism before objecting, because the owner is right that both conventions exist. The difference here is not convention, it is that units_blocked grew during the measurement window, so the capacity-based number rises partly because supply was withdrawn rather than because more nights were sold.
- Separate the two movements numerically: hold units_blocked at its pre-period level and recompute the capacity-based figure. Whatever remains of the 2.1pp after that is the real effect, and the difference is the withdrawal.
- Reframe the blocking as a finding rather than an inconvenience. Hosts blocking dates during a launch is a supplier-side signal the launch owner needs, and it may matter more than the occupancy delta.
- Put both numbers in the document with their denominators named in the row labels, plus the blocked-nights level as a separate line, so the reader can see the mechanism without being told which number to prefer.
- Make the ask specific and small: pick one convention for this review, state it in the chart title, and never mix the two inside a single comparison. This is easier to agree to than a debate about which figure is right.
Follow-up
- The owner says the blocking is seasonal and unrelated to the launch. How do you test that?
- Your manager wants you to let it go because the difference is small. What do you do?
- How would you stop this ambiguity reaching a review document next time?
- 01
Tell me about a time you disagreed with a product manager’s decision; how did you handle it?
- 02
A supplier-growth team moved one market's inventory from 'strict' to 'flexible' cancellation_policy in fct_rate_availability_snapshot. Their dashboard shows bookings up 14% over eight weeks. Your read on fct_stay_night shows stayed nights flat, and traveller-initiated cancellation up from 18% to 31% in the same lead-time bucket. Their quarterly goal is booked nights, and the lead has already sent the 14% to their director. You have ten minutes in their weekly review. Prepare what you open with, what you concede, and what you will not soften.
- 03
A launch review is tomorrow. Occupancy computed on sellable nights, SUM(units_sold_to_date + units_available) from fct_rate_availability_snapshot, shows the launch market up 0.4pp. Computed on physical capacity, capacity_units from dim_supply_unit, it shows up 2.1pp, because individual_host suppliers blocked dates during the period and units_blocked is excluded from the first denominator but inside the second. The launch owner asks you to use the second, noting it is a documented convention. Prepare your response and what you put in the review document.
Is this an official Priceline interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Priceline. Rounds and questions reflect what candidates have reported, not a process Priceline has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How much time should I spend preparing for the take-home assignment?
Treat the take-home assignment as a high-quality deliverable. While there is no fixed time limit, ensure your code is clean, well-documented, and your conclusions are clearly supported by the data you analyzed.
PracHub interview research ↗Is the interview process mostly technical or behavioral?
It is a balanced mix. You will have dedicated technical rounds, but the final round heavily weighs how you approach ambiguous business problems and communicate your reasoning.
PracHub interview research ↗What is the best way to stand out during the interview?
Show that you understand the business. The best candidates don't just solve the math; they explain how their solution will help Priceline improve the user experience or business performance.
PracHub interview research ↗Does Priceline value specific programming languages?
Proficiency in Python or R is expected for data analysis, but SQL is the bedrock of the role. You must be able to write complex queries fluently.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22