As a Data Scientist at Tripadvisor, you will play a pivotal role in harnessing data to enhance the user experience and drive business decisions. This position is crucial for shaping the way users interact with various products, from personalized recommendations to optimized search results. Your work will directly influence how millions of travelers access and utilize information, making their planning and booking processes more efficient and enjoyable.
In this role, you will collaborate with product managers, engineers, and other data scientists to turn data insights into actionable strategies. You will engage with complex datasets to develop machine learning models that optimize recommendations, improve user engagement, and support various business initiatives. The scale and diversity of data at Tripadvisor present unique challenges and opportunities, making this role both critical and intellectually stimulating. You will contribute to projects that have a significant impact, such as enhancing the accuracy of pricing models and improving the relevance of search results for users worldwide.
Screening Interview
reportedAn extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.
What to demonstrate
- Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
- Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
- How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one
How to prepare
- Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
- Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
- Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
Case Study Interviews
reportedThis round runs as a working session, so part of what it decides is whether you are useful to think with. The interviewer will interrupt: a hint that the data you assumed does not exist, a challenge to your metric, a nudge toward a branch you skipped. Treating those as interference is the common failure. Reason out loud while your thinking is still provisional so there is something to react to, and when a redirect arrives, take it instead of defending the path you had already started down.
What to demonstrate
- Whether your reasoning is audible while it is still unsettled, or only after you have privately decided
- What you do with a hint: absorb it and adjust, or argue for the original route
- Whether your clarifying questions have answers that would change your approach, as opposed to filling silence
- Whether you can be wrong about something in the middle of the case and keep moving without restarting
How to prepare
- Run practice cases with a partner instructed to interrupt twice: once to remove a data source you assumed existed, once to reject the metric you chose. Practise absorbing both without going back to the start.
- Before each practice case, write down the clarifying questions you plan to ask, then check afterwards whether any answer actually changed what you did. Drop the ones that did not.
- Explain an analysis you already know well to someone outside the field and have them stop you at every point where the reasoning jumped a step.
Final Interviews
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub editorial advice for the preparation topics above.
Reading cancellation, completion or repeat rates on cohorts that have not matured
A cohort of bookings made last week for stays six months out cannot have cancelled at the check-in gate yet, so its cancellation rate is mechanically near zero and its completion rate mechanically near zero as well, in opposite directions. Comparing that cohort with a mature one is not a noisy comparison, it is a guaranteed wrong one, and the bias always makes the recent period look different in a way that invites a false story about a recent change. Because lead time is heavily right-skewed, the mean lead time is a bad maturity threshold; use the cohort's 95th percentile, or report a hazard at a fixed age (cancelled within k days of booking) with k capped at the youngest cohort's elapsed age. The same applies to repeat rate, where the honest answer is often that the cohort in question is not readable for another nine months.
Running per-user experiments on shared, finite inventory
Independence between units fails when they compete for the same rooms on the same dates: a unit booked by a treatment user is removed from the control users' result sets, so the control arm is degraded by the treatment and the measured lift is inflated, sometimes to the point where a neutral change reads as a clear win. The effect is largest exactly where it matters most, in constrained markets and peak dates, and it is invisible in the usual diagnostics because both arms look balanced on pre-period covariates. Supplier-side treatments have the same problem in reverse, since a pricing or ranking change applied to some inventory changes the alternatives that every traveller sees. Randomise at the level that contains the competition, which is normally market-by-stay-week or a switchback on a market, and power the test on the number of clusters, not the number of users.
Naming a model class before naming the deployment constraints
Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.
Averaging per-user rates to produce a population rate
Decide which quantity you want: the mean of per-user ratios and the ratio of summed numerator to summed denominator are different estimands, and heavy users dominate one but not the other. For a ratio metric, aggregate numerator and denominator separately and use the delta method for its variance.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
If tasked with improving a recommendation algorithm, what steps would …
If tasked with improving a recommendation algorithm, what steps would you take?
Approach
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Write invariant checks for a rate availability snapshot
The DataFrame snap holds one row per supply_unit_id per forward stay_date per snapshot_date, with units_available, units_sold_to_date, units_blocked, listed_price_usd (nullable), is_closed_to_arrival, min_length_of_stay and lead_time_days. Write check(snap) returning a tidy failure frame with columns check_name, supply_unit_id, stay_date, snapshot_date, detail, plus a per-check summary count. Cover at least: primary-key uniqueness, lead_time_days equal to (stay_date - snapshot_date) in days, listed_price_usd null exactly when units_available is 0, and non-negative counts. It must return the right empty frame on empty input.
Approach
- Write each check as a boolean mask over the whole frame, not a row-wise function, and append a small frame carrying check_name plus the three key columns and a
detailstring built from the offending values. Concatenate at the end with a typed empty frame in the list so an all-clean or empty input still returns the declared columns and dtypes. Declare the check names once as a module-level tuple, CHECK_NAMES, and drive both the run loop and the summary off that tuple rather than off whatever happened to fail. - Test nullity with .isna(); listed_price_usd is a float column so a NULL is NaN, and NaN == NaN is False, which means any
== Noneor== np.nancomparison returns an all-False mask and the check passes vacuously on every row. - Check the price rule in both directions: a non-null price where units_available is 0 means someone imputed the last known price onto a sold-out unit-date, which is the single most damaging corruption here because it turns censored observations into apparently-observed ones for any elasticity work.
- Separate hard invariants from business expectations and label them with a severity column. Uniqueness of (supply_unit_id, stay_date, snapshot_date), the lead-time identity and non-negativity are invariants. Monotonicity of units_sold_to_date across snapshot_date is not: cancellations legitimately reduce cumulative pickup, so it belongs at warn severity with a threshold on the size of the drop.
- Build the summary by grouping the failure frame on check_name and then reindexing against CHECK_NAMES with fill_value=0. A bare groupby emits no row for a check that found nothing, so a clean input returns a summary listing only the broken checks and an empty input returns no summary at all — which on a dashboard is indistinguishable from the suite never having run, the one state a summary exists to rule out. Return counts per stay_date alongside it so a caller can see whether failures are uniform or concentrated on a single partition, which is what distinguishes a code bug from one bad ingest day.
Worked solution 20 min
- Declare CHECK_NAMES as a fixed tuple of the four check names, then define a helper that takes a mask and a check_name and returns the key columns of the failing rows plus a detail string; collect results in a list.
- Check 1: snap.duplicated(subset=['supply_unit_id','stay_date','snapshot_date'], keep=False) flags every member of a duplicate key group, not just the later ones.
- Check 2: (snap['stay_date'] - snap['snapshot_date']).dt.days != snap['lead_time_days'].
- Check 3: two masks, (units_available == 0) & listed_price_usd.notna(), and (units_available > 0) & listed_price_usd.isna().
- Check 4: any of units_available, units_sold_to_date, units_blocked below zero; and min_length_of_stay below 1.
- Concatenate with an explicit empty typed frame, then build the summary as failures.groupby('check_name').size().reindex(CHECK_NAMES, fill_value=0), so a check that found nothing reports 0 instead of vanishing and a clean or empty input still returns one row per declared check.
Follow-up
- Which of your checks would you make a hard pipeline failure and which a warning, and what is the cost of getting that wrong in each direction?
- units_sold_to_date drops by 4 between two consecutive snapshots for one unit-date. Give two innocent explanations and one that should page someone.
- How would you run these on a table too large to hold in memory, keeping the same failure frame?
Permutation test occupancy on market-week randomised clusters
A ranking change was randomised over 48 clusters, each a destination market crossed with a stay week, 24 per arm. You get clusters with destination_market_id, stay_week, arm, stayed_nights, sellable_nights; markets recur across several weeks. The metric is occupancy computed as a ratio of sums, treatment minus control. Write a permutation test from scratch with 10,000 reshuffles and a bootstrap confidence interval, using no scipy or statsmodels testing function. Report the observed difference, a two-sided p-value and a 95% interval, and state the resampling unit you chose for each.
Approach
- Compute the observed statistic as a ratio of sums within each arm, not a mean of per-cluster occupancies, so a market-week with 9,000 sellable nights does not carry the same weight as one with 300.
- For the permutation null, shuffle the arm label vector across clusters. Randomisation was independent per market-week, so the cluster is the exchangeable unit; permuting the night-level rows instead destroys the cluster correlation and shrinks the null distribution by roughly the square root of the nights per cluster, which turns almost any observed difference into a significant one.
- Vectorise the reshuffles: generate a (10000, 48) matrix of random values, argsort each row, and use the first 24 positions as the treatment index set, so the whole null distribution is built with array operations rather than a Python loop over 10,000 iterations.
- Use the add-one p-value, (1 + count of |permuted| >= |observed|) / (B + 1). The uncorrected version can report exactly zero, which claims more certainty than 10,000 reshuffles can support.
- For the interval, resample whole markets with replacement rather than clusters, because a market's weeks share demand conditions and are not independent draws; then recompute the ratio-of-sums difference on each resample. Report the implied minimum detectable effect at 48 clusters alongside the p-value, so a null result is read as underpowered rather than as evidence of no effect.
Follow-up
- Your interval contains zero at 48 clusters. What cluster count would you need for an 80% chance of detecting a 2 percentage point occupancy lift, and what does that cost in calendar time?
- Why is a per-traveller randomisation on this same change biased, and in which direction?
- Which guardrails would you read alongside occupancy, and which failure mode does each one catch?
Can you explain the time complexity of your code?
Can you explain the time complexity of your code?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Write a SQL query to extract specific data from a given database.
Write a SQL query to extract specific data from a given database.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Find sold-out runs of three or more consecutive nights
From fct_rate_availability_snapshot (supply_unit_id, stay_date, snapshot_date, units_available, units_blocked, is_closed_to_arrival), take the last snapshot on or before each stay_date and find every run of three or more consecutive stay_dates where a unit was genuinely sold out: units_available = 0 and units_blocked = 0. Return supply_unit_id, run_start, run_end and run_length_nights for one destination market over the next 180 days. A stay_date with no snapshot row must break the run rather than bridge it.
Approach
- Build the as-of picture first, one row per (supply_unit_id, stay_date) from the latest snapshot_date on or before that stay_date, so the run is defined over a single value per unit-date instead of over every capture.
- An absent stay_date breaks the run by itself under the island key below, so no generated date spine is needed to stop bridging. Across two consecutive sold-out dates, stay_date and the row number each advance by one and stay_date minus row number holds constant; across a hole of g days, stay_date advances by g+1 while the row number advances by 1, so the key jumps by g and a new island starts. Sold-out dates of 1, 2, 3, 5, 6 and 7 January give keys 31 Dec, 31 Dec, 31 Dec, 1 Jan, 1 Jan, 1 Jan, which is two islands with or without a spine.
- Build a spine only to measure coverage, not to repair the runs. This query cannot distinguish a stay_date with no snapshot from one whose snapshot showed units free, because both are simply absent from the filtered set; if the reader needs that distinction, report missing unit-dates as their own column rather than folding them into the run logic.
- Keep only rows with units_available = 0 AND units_blocked = 0. An owner-blocked or out-of-order date was never open for sale, and counting it as sold out overstates compression to whoever reads the list.
- Number surviving rows per unit with ROW_NUMBER() OVER (PARTITION BY supply_unit_id ORDER BY stay_date) and group on (stay_date minus rn days), which is constant inside a consecutive run and changes across a break. The row number must be computed after the sold-out filter, not before it.
- Aggregate per island with MIN and MAX stay_date and COUNT(*) as the length, filter to length >= 3, and order by length descending.
Worked solution 40 min
- Build and materialise the as-of table, then compare its row count against units multiplied by 180 days. The shortfall is the number of unit-dates carrying no snapshot at all, and each one will break a run rather than bridge it.
- Apply the sold-out filter and record how many unit-dates survive; compare against the count when units_blocked is ignored, and quote the difference as the size of the blocked-capacity error.
- Add ROW_NUMBER and the island key, then verify by hand on one unit that deleting a middle date splits its run into two whose lengths sum to one less than the original. Verify as well that adding a date spine changes neither the runs nor the counts, since a NULL-availability row fails the sold-out filter anyway.
- Aggregate to islands, assert the length identity, and apply the three-night filter last.
Follow-up
- Extend this to market-level compression: how many units in a market are sold out on the same night, and what threshold would you alert supply acquisition on?
- Some of these runs are minimum-length-of-stay restrictions rather than genuine sell-outs. How do you tell them apart with the columns in this table?
- One run spans a moving holiday that shifted two weeks year over year. How do you compare it with the same period last year?
How would you evaluate the success of a new feature launched on the pl…
How would you evaluate the success of a new feature launched on the platform?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you set pricing for a new listing with various constraints?
How would you set pricing for a new listing with various constraints?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How do you optimize your code for performance?
How do you optimize your code for performance?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What metrics would you use to evaluate a recommendation system?
What metrics would you use to evaluate a recommendation system?
Approach
- Fix the population and the time window before naming any metric.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Explain the concept of A/B testing and how you would implement it.
Explain the concept of A/B testing and how you would implement it.
Approach
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
- Say whether units interfere with each other, and switch design if they do.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
What is the difference between supervised and unsupervised learning?
What is the difference between supervised and unsupervised learning?
Approach
- Work from the decision backwards to the evidence you would need.
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Define success for a flexible-date search results page
A flexible-date results view lets a traveller shift check-in by up to three days to find a cheaper night. You have fct_search (trip_intent_key, visitor_id, traveller_id, searched_at_utc, destination_market_id, check_in_date, length_of_stay_nights, lead_time_days, results_returned_count, is_bot_flagged), fct_booking (trip_intent_key, booked_at_utc, booking_status, gross_booking_value) and fct_stay_night (night_status, unit_revenue_usd). Define success: give the metric tree from search to completed stay, one primary metric stated as numerator, denominator, window and date axis, and two guardrails. Say which grain the denominator uses and why.
Approach
- Fix the demand unit before anything else: collapse fct_search on trip_intent_key, because this feature directly changes how many search rows one intent produces, so a search-grain denominator tracks the interface rather than demand.
- Lay the tree in stages that each have a real denominator: intents with results_returned_count > 0, then intents with a booking within 7 days of first search, then bookings reaching night_status = 'stayed', then stayed nights and unit_revenue_usd.
- Pick the primary as the qualified intent-to-booking rate on the booking axis, cohorted on the first search date for the key and held 7 days for the attribution window to close, because the feature acts at search time while the stay outcome is months away.
- Choose guardrails that catch what this feature can quietly trade away: nights per booking and ADR on stayed nights at a fixed reference exchange rate, since shifting dates can sell cheaper and shorter trips, plus traveller-initiated cancellation rate by lead-time bucket.
- Pre-register the true readout on the stay axis - completed stay-nights and contribution margin per completed stay-night - with the date it becomes readable, and keep searches per intent as a published diagnostic outside the rate.
Worked solution 20 min
- Write the denominator first: DISTINCT trip_intent_key with is_bot_flagged = FALSE and results_returned_count > 0, cohorted on the minimum searched_at_utc per key.
- Write the numerator as those keys with a matching fct_booking row booked within 7 days of that first search, counting the booking whether or not it later cancels.
- Extend the tree forward with left joins to fct_stay_night so lost stages stay visible as NULLs, and record the row count at each grain before aggregating.
- State the splits that must accompany the rate: destination_market_id, device_type and lead-time bucket, since flexible dates behave differently at 3 days out and at 200.
- Write the guardrail definitions with the same rigour as the primary, including the fixed reference exchange rate for ADR.
Follow-up
- Searches per intent falls 20 percent and the intent-to-booking rate is flat. Has the feature worked?
- How would you separate a traveller who genuinely moved their dates from one who would have booked the same date anyway?
- What breaks if you build this funnel by joining on traveller_id instead of trip_intent_key?
Stay-nights fell nine percent against a shifted holiday week
Completed stay-nights in destination market 412 fell 9% week over week, and 6% against the same ISO week last year. You have fct_stay_night (stay_date, night_status, is_weekend, market_holiday_flag, destination_market_id) and fct_rate_availability_snapshot for sellable nights. Nothing shipped that week and no pipeline alert fired. Establish whether the fall is demand or calendar composition, and hand back a like-for-like number the commercial team can act on. Name the date axis you used, and the occupancy denominator if you quote one.
Approach
- Break the weekly total into a daily series first. A week is seven numbers, and a 9% weekly gap is usually one or two dates carrying almost all of it, which immediately rules out a broad demand story.
- Attach is_weekend and market_holiday_flag from the market's local calendar to every stay_date, then count day types per week. A week holding one fewer Friday-Saturday pair loses nights at the highest-occupancy, highest-ADR day type with nothing behavioural having changed.
- Rebuild the comparison by direct standardisation: compute mean stayed nights per day type in each week, then reweight the comparison week to the base week's day-type composition. Roll rates up by re-summing numerator and denominator, never by averaging daily rates.
- For the year-over-year cut, align on market_holiday_flag rather than on ISO week number. School terms and moving holidays shift by weeks, so the same calendar week is not the same demand week, and a date offset does not fix it.
- Check the denominator moved with the numerator. If sellable nights fell by a similar proportion, this is capacity leaving the calendar, not demand leaving the market, and the answer goes to a different team.
Follow-up
- The holiday and the missing weekend night explain six of the nine points. How do you present the residual three so nobody reads it as the start of a trend?
- This market's holiday moved but the neighbouring market's did not. What does comparing them buy you, and what does it not control for?
- You have to produce this comparison automatically every week for 300 markets. What do you precompute, and where does the automation get it wrong?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Work that nobody used is a common and unflattering pattern in data careers, and interviewers probe for it. Have a story about an analysis that changed a decision, and be specific about how you got it in front of the person who could act. Also have one about work that went nowhere, with your reading of why.
Tell me about a time you faced disagreement within your team. How did …
Tell me about a time you faced disagreement within your team. How did you handle it?
Approach
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
State your own impact when the evidence is confounded
You built the pickup-based stay-week forecast that replaced a trend model, and it was adopted across markets last quarter. Forecast error on stay-week stayed nights improved 9% over the quarter. In the same quarter two markets were added to the portfolio, the seasonality baseline was re-fitted by another team, and one large market's moving holiday realigned against the prior year. You are writing the impact section of your own review. Write the claim you would actually make, the evidence behind it, and what you explicitly do not claim.
Approach
- Name the counterfactual before the number: 9% against what the old model would have produced on the same weeks, not against last quarter's realised error, since the portfolio and the calendar both changed underneath.
- Build the honest comparison by back-testing the old model on the current quarter's stay weeks and scoring both models on the same weeks, same markets, same horizon. This removes the two added markets and the re-fitted baseline from the comparison entirely, because both models see the same inputs.
- Handle the moving holiday by excluding or separately reporting the affected stay weeks, since a realignment changes the difficulty of the forecast rather than the quality of either model, and it would otherwise sit inside your credit.
- Split the claim into what the method did and what you did. Adoption across markets involved other people; the defensible personal claim is the method and the migration you drove, with the organisational result reported alongside and attributed plainly.
- Write the non-claim explicitly and early. Stating that you cannot attribute the portfolio-level improvement to the model alone is what makes the part you do claim credible, and it pre-empts the reviewer finding the confound themselves.
Follow-up
- The head-to-head back-test shows only a 3% improvement. What do you write?
- How would you have instrumented the rollout so this was measurable without a back-test?
- A peer's review claims the same 9%. How do you handle that?
Retract an occupancy comparison after the decision shipped
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
Approach
- Quantify the error before telling anyone, because the first question will be 'how wrong'. Recompute both markets on sellable nights, which excludes units_blocked, and report the corrected gap and its sign, not just that the original was wrong.
- Separate the numerical error from the decision error. The comparison was invalid, but the spend move may still have been right; establishing whether the decision would have flipped is a different analysis and it is the one the business needs.
- Tell the decision-maker directly and first, before it appears in a dashboard or a peer surfaces it. Order matters because being told by a third party converts a mistake into a credibility problem.
- Write the note with the correction, the size, the decision implication, and the reversal cost in that order. Three weeks of moved spend has a real cost to undo, and a correction that does not price the reversal forces the reader to do the work you skipped.
- Name the specific control that would have caught it and put it in place in the same note: an assertion in the query that both arms use the same denominator expression, and the denominator named in the chart title. A retraction without a mechanism reads as an apology rather than a fix.
Follow-up
- The corrected numbers still support the original decision. Do you still send the note, and does it read differently?
- Your manager suggests quietly fixing the dashboard and not raising it. How do you respond?
- What would you have had to do differently three weeks ago, in the query itself, rather than in your review habits?
- 01
Tell me about a time you faced disagreement within your team. How did you handle it?
- 02
You built the pickup-based stay-week forecast that replaced a trend model, and it was adopted across markets last quarter. Forecast error on stay-week stayed nights improved 9% over the quarter. In the same quarter two markets were added to the portfolio, the seasonality baseline was re-fitted by another team, and one large market's moving holiday realigned against the prior year. You are writing the impact section of your own review. Write the claim you would actually make, the evidence behind it, and what you explicitly do not claim.
- 03
Three weeks ago you published a market occupancy comparison that led to marketing spend being moved out of one market. Reviewing the query, you find the two markets used different denominators: sellable nights from fct_rate_availability_snapshot (units_sold_to_date + units_available) for one, and capacity_units from dim_supply_unit for the other. The second market is mostly individual_host supply with large units_blocked, so its occupancy was understated. The spend move is already live. Prepare the correction: who you tell, in what order, and what the note says.
Is this an official Tripadvisor interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Tripadvisor. Rounds and questions reflect what candidates have reported, not a process Tripadvisor has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the interviews, and how much preparation time should I expect?
Interviews at Tripadvisor can be challenging, particularly in technical areas. Candidates often report spending several weeks preparing, focusing on both technical skills and behavioral questions.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates often exhibit a strong understanding of data science principles, effective problem-solving abilities, and a collaborative spirit. Demonstrating these qualities will set you apart.
PracHub interview research ↗What is the culture like at Tripadvisor?
Tripadvisor promotes a collaborative and innovative work environment. Team members are encouraged to share ideas and work together to enhance the user experience.
PracHub interview research ↗What is the typical timeline from the initial screen to an offer?
The timeline can vary, but candidates often receive feedback within a couple of weeks after interviews. Expect multiple rounds of interviews, especially for technical roles.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22