Airbnb · Data Scientist
Updated · 2026-09-22

Airbnb Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

This guide covers what a Data Scientist at Airbnb is expected to do and how to prepare for the interview.

A large share of questions open as "how would you measure X", where the real work is choosing the metric, fixing its denominator, and defining the population it applies to. Any computation comes last and is frequently not required at all.

Airbnb candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Design switchback tests that survive supply-side interferenceEstimate cross-side elasticities from cohort and holdout dataDecompose completed orders into requests, fill, completion

32 min read

Practice 17 Data Scientist prompts
24Company bank questionsSnapshot · Sep 22, 2026 PT
25Candidate experiences ↗Read their reports
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

This guide covers what a Data Scientist at Airbnb is expected to do and how to prepare for the interview.

01

Recruiter Screening Call

reported

Most candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.

What to demonstrate

  • Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
  • Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
  • The substance of the questions you ask back, which an experienced screener reads as a level signal

How to prepare

  • Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
  • Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
  • Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
PracHub interview research
02

Take-Home Assignment

reported

Your submission is read asynchronously by someone who cannot ask you a clarifying question, and who will usually skim it before reading it properly. That changes what good looks like. The conclusion belongs near the top with the supporting analysis beneath it, and the code should run start to finish on a clean machine without a manual step you forgot to document. A reviewer forced to reconstruct your reasoning from the order of notebook cells is already discounting the work. The submissions that land are the ones where a busy reader gets the answer immediately and can verify it if they want to.

What to demonstrate

  • Whether the answer arrives before the methodology, so a reader who stops after the first page still has the recommendation
  • Whether the code runs end to end from the submitted files, with dependencies and data paths declared rather than assumed
  • Whether each chart is legible on its own, carrying axis labels and units, and exists to support a claim made in the text

How to prepare

  • Restructure a past analysis so the opening paragraph holds the recommendation and the number behind it, then check that nothing later in the document quietly contradicts it
  • Copy your own submission into an empty directory, run it in a clean environment, and fix everything that breaks; hidden local state is caught here or by the reviewer
  • For every chart, write the one sentence it is meant to prove, and delete the chart if you cannot write that sentence
PracHub interview research
03

Technical Coding Screen

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research
04

Virtual or On-Site Loop

reported

A day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.

What to demonstrate

  • Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
  • Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
  • Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method

How to prepare

  • Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
  • Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
  • Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
  • Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
PracHub interview research

25 candidate reports. Individual accounts describe a particular role and hiring cycle.

Machine Learning Engineer

Airbnb Senior+ Machine Learning Engineer Interview Experience — Search Ranking and Team Matching

Technical Screen → Onsite

A referred applicant to Airbnb’s personalization and relevance organization reports recruiter contact after two weeks. An initial coding screen asked for book titles to be formatted as a table. Four days later, the applicant advanced to a virtual onsite covering document redaction, project decisions, search ranking without text queries, and retrieval-assisted guidance for customer-support staff.…

Read full experience
Data Engineer

Airbnb Data Engineer Interview Experience — Four-Round Onsite, Then a Recruiter Who Ghosted Me

Online Assessment → OnsiteOutcome: rejected

Not many interview reports out there, so let me add mine — hope it helps. My recruiter was unreliable and ghosted me right after the interview. One round of OA. The VO (onsite) was four rounds of analytics interviews. OA: The SQL was pretty easy. The Python question gave a list/dict of city housing prices, and I had to find the median price for each city. VO: SQL: Find the top N, then figure out…

Read full experience
Software Engineer

Airbnb Software Engineer Interview Experience — A Perfect Cloud-Storage OA Score Followed by Rejection

Online AssessmentOutcome: rejected

I received the OA immediately after applying, got a perfect score, then received a rejection email a week later. The problem was a cloud storage system. It had four levels. You only got the next level after finishing and submitting the current one. Level 1: Support adding new files, retrieving files, and copying files. Level 2: Support finding files by matching prefixes and suffixes. Level 3: Sup…

Read full experience
Software Engineer

Airbnb Software Engineer interview experience

Online Assessment

I applied for a Staff Software Engineer role and then heard nothing for more than a month. When I finally got an email, it invited me to a CodeSignal technical interview. The interface looked polished and had practice exercises, but the editor felt limited compared with a VS Code-style workspace and lacked conveniences such as autocompletion. The HackerRank-style assessment allowed 90 minutes and…

Read full experience
Software Engineer

Airbnb Software Engineer Interview Experience — A Four-Level Banking System Online Assessment

Online Assessment

I recently took Airbnb's software engineering online assessment. I hope this question format helps people who take it later. One of the online assessment questions was a banking system or filesystem with unit tests. It was similar to CodeSignal's progressive-level problems. It was not a set of independent LeetCode questions; instead, new requirements kept being added to the same codebase. The ove…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Conditioning the analysis on completed orders

Wait-time distributions, price elasticities and rating models fit only on completed orders are conditioned on an outcome that the intervention itself changes. The requests that never matched, or that the consumer abandoned, are the population a liquidity fix targets, so excluding them biases every estimate toward the status quo and can flip the sign of a price elasticity. Any query starting FROM fct_order is already inside this trap; start from fct_request and left join.

02

Treating abandoned requests as missing rather than censored

Consumers who give up before matching are censored observations, so the mean time-to-match computed over matched requests understates true waiting and improves mechanically whenever abandonment rises. A change that makes people quit sooner will look like a latency win. Use survival methods with abandonment as the censoring event, or never report time-to-match without reporting abandonment next to it.

03

SQL that silently fans out on a one-to-many join

State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.

04

Reaching for a model before the target metric exists

Before naming an algorithm, write down the label, the prediction time, and the action that changes when the score crosses a threshold. If you cannot say what decision the output drives, any modelling choice is guesswork dressed up as method.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

How do you calculate confidence intervals for ratios or non-linear met…

medium
statistics and probability

How do you calculate confidence intervals for ratios or non-linear metrics derived from skewed datasets?

Approach
  1. Say what the estimate is of, and over what population it generalises.
  2. Sanity-check the answer against a simple bound or a simulated case.
  3. Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
  • How would you explain this result to someone who does not know statistics?
  • What sample size would you need to detect an effect half this size?

When is propensity score matching preferable to regression adjustment …

medium
machine learning and modelling

When is propensity score matching preferable to regression adjustment in estimating the causal impact of a customer support intervention?

Approach
  1. Check what information would not exist at prediction time, and exclude it.
  2. Set a baseline first, so any model has something honest to beat.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • Where could label leakage enter this setup?

Simulate dispatch cascades and the censoring of wait time

mediumWorked solution
simulationcensoringmonte carlo error

A request is offered to one provider at a time. Each offer resolves 12 seconds after it is sent, and each provider accepts independently with probability 0.55. After six declines the request is marked no_supply. Independently, the consumer abandons at time A drawn from an Exponential distribution with mean 90 seconds; abandonment before a pending offer resolves ends the request unmatched. Simulate 200,000 requests and report: the share matched, the mean time-to-match over matched requests, and the mean over requests that would have matched with abandonment switched off. Give a Monte Carlo standard error for the share.

Approach
  1. Vectorise the cascade: draw K with np.random.default_rng().geometric(0.55), mark K > 6 as no_supply, and draw A = rng.exponential(90) independently; the match condition is K <= 6 and A > 12*K. Looping request by request is the difference between two seconds and two minutes of runtime.
  2. Compute both means on the same draws so the comparison is paired and the difference is not itself a Monte Carlo artefact.
  3. Recognise the structure driving the answer: abandonment censors long cascades harder than short ones, so conditioning on matched requests is not a neutral filter, it is a filter correlated with the quantity being measured.
  4. Quote the share to three decimals only: the standard error of a proportion is sqrt(p(1-p)/n), roughly 0.0009 at n = 200,000, so further digits are noise.
  5. Check against the closed form P(match) = sum over k of 0.45^(k-1) * 0.55 * exp(-12k/90) for k = 1..6; a simulation with no analytical check is an untested function.
Worked solution 30 min
  1. k = rng.geometric(0.55, size=200_000); a = rng.exponential(90.0, size=200_000); t = 12.0 * k.
  2. supplied = k <= 6; matched = supplied & (a > t).
  3. share = matched.mean(); se = sqrt(share * (1 - share) / 200_000).
  4. observed_mean = t[matched].mean(); latent_mean = t[supplied].mean().
  5. Compare share against the closed form 0.7911 and print the gap in standard errors.
EXPECTED RESULTMatched share about 0.791 (closed form 0.7911, Monte Carlo SE about 0.0009); mean time-to-match over matched requests about 19.5 seconds; latent mean over all requests with K <= 6 about 21.2 seconds; the observed mean therefore understates the true wait by roughly 1.7 seconds, about 8%.
Follow-up
  • A change ships that makes consumers abandon sooner. What happens to your reported mean time-to-match, and how would you report latency so that this cannot look like a win?
  • How would you estimate the same quantity from production data, where you never observe the latent match time of an abandoned request?

Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Diagnostic, scored before you study anything
  • Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
  • Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
  • Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.

Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Largest gap: find the boundary rather than the subject
  • Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
  • Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
  • Re-solve one of them from memory four hours later, on paper, with nothing open.

Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Largest gap: drill the blocking sub-skill
  • Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
  • Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
  • Have someone else read your one-sentence rule and find the precondition you left out.

Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Second gap, plus maintenance on your strongest area
  • Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
  • Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
  • Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.

Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05The gap that is not a skill
  • Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
  • Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
  • Deliver one answer again with a hard 90-second limit to force structure before detail.

Deliverable: Two recordings with a counted improvement in time-to-first-question.

Practice prompt ↗Practice prompt ↗
06Retest under day-one conditions
  • Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
  • Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
  • Write which single block you would still lose the offer on.

Deliverable: A second scored rubric placed next to the first, with one named remaining risk.

Practice prompt ↗Practice prompt ↗
07Full loop under interview conditions
  • Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
  • Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
  • Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.

Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Interviewers here are not checking whether you can describe a project. They want the decision you made, why you made it under the information you had, and what changed afterwards that someone else could measure. A story that ends at 'I built a model' has no ending. Say what the model caused, or what you stopped doing because of it.

How do you prioritize competing requests from product, engineering, an…

medium
behavioural and stakeholder questions

How do you prioritize competing requests from product, engineering, and operations stakeholders when your team has limited bandwidth?

Approach
  1. Quantify the outcome, including what you would not claim credit for.
  2. Name the disagreement or constraint, and how you resolved it with evidence.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What would you do differently if you ran that project again?
  • What did you decide not to do, and why?

Explain a switchback confidence interval to a non-technical executive

easy
communicating uncertaintyswitchbackdecision framing

A switchback test of a dispatch-radius change ran 1,152 market-hour blocks across six markets. SLA fill rate moved +1.8 percentage points, 95% interval [-0.4, +4.0], variance clustered at the block. Those markets serve about 250,000 eligible requests a week at 88% fill and 93% completion. An executive with no statistics background wants a ship-or-wait answer inside a five-minute update. Deliverable: the two-minute spoken explanation, your recommendation, and the single condition that would change it. You may not use the words significant, p-value, or confidence interval.

Approach
  1. Open with the decision and the recommendation, then justify; an executive who hears the caveat first stops listening before the ask arrives.
  2. Translate both interval bounds into the unit the executive already manages: eligible requests times percentage points times completion rate gives weekly completed orders, so the range becomes 'between about 1,000 fewer and about 9,300 more completed orders a week, best single guess about 4,200 more'.
  3. Say plainly what the range does and does not rule out: it does not rule out a small loss, and it is wide because the test has 1,152 effective units, not 250,000 consumers. Block-level randomisation is the reason the sample is small, and it is the reason the number is trustworthy at market level.
  4. Price the two errors against each other: a reversible dispatch parameter with a bounded downside is cheap to ship and cheap to revert, so the decision rule is not 'is the effect proven' but 'is the worst case affordable and detectable'.
  5. End with the one condition that flips you: name the monitoring metric (provider utilisation and idle time, since a wider radius can raise fill by burning provider hours) and the threshold at which you revert.
Follow-up
  • How many more weeks of blocks would it take to halve the width of that range, and is that worth the delay?
  • The executive asks 'so is it real or not' - what do you say without reaching for statistical vocabulary?
  • What would you monitor post-ship that the experiment itself could not measure?

Turn a one-line supply request into a scoped analysis

easy
scopingliquidity diagnosismetric definition

A cross-functional lead messages: 'can you look into whether we have enough drivers in the north zone?' No metric, no deadline, no stated decision. You have fct_supply_session (online_seconds, engaged_seconds, idle_seconds, market_id, offline_at_utc), fct_request (request_status, requested_at_local, market_id, dispatch_attempts) and dim_market (timezone). Deliverable: the three questions you ask before touching data, then a scoping note under 150 words naming the decision it serves, the one number that answers it with numerator, denominator and window, the first cut you run, and the date you return.

Approach
  1. Ask what lever is being considered, because 'enough supply' has no definition independent of the action: a provider incentive budget, a dispatch radius change and a recruiting target each need a different number and a different window.
  2. Ask what happens if the answer is 'yes, enough', since a question with the same action under both answers is not worth running, and ask when the decision is made, which sets the depth you can afford.
  3. Translate the vague word into a measurable pair with opposite signatures: share of requests with request_status = 'no_supply' plus provider utilisation (SUM(engaged_seconds) / SUM(online_seconds)) by market-hour. High utilisation with rising no_supply is supply-constrained; high idle_seconds with flat request counts is demand-constrained.
  4. State the data handling that changes the answer: clip open sessions (offline_at_utc IS NULL) at the window edge rather than dropping them, and cut by requested_at_local using dim_market.timezone, because supply shortage is an hour-of-day phenomenon that a UTC cut smears away.
  5. Commit to a first cut and a return date, and name explicitly what you are not doing, so the scope can be argued with before the work rather than after it.
Follow-up
  • The lead says 'just give me the driver count' - how do you respond without refusing the request?
  • What single additional table would let you distinguish a dispatch problem from a genuine supply shortage?
  • How does your answer change if the zone is one part of a market rather than a market_id of its own?
  • 01

    How do you prioritize competing requests from product, engineering, and operations stakeholders when your team has limited bandwidth?

  • 02

    A switchback test of a dispatch-radius change ran 1,152 market-hour blocks across six markets. SLA fill rate moved +1.8 percentage points, 95% interval [-0.4, +4.0], variance clustered at the block. Those markets serve about 250,000 eligible requests a week at 88% fill and 93% completion. An executive with no statistics background wants a ship-or-wait answer inside a five-minute update. Deliverable: the two-minute spoken explanation, your recommendation, and the single condition that would change it. You may not use the words significant, p-value, or confidence interval.

  • 03

    A cross-functional lead messages: 'can you look into whether we have enough drivers in the north zone?' No metric, no deadline, no stated decision. You have fct_supply_session (online_seconds, engaged_seconds, idle_seconds, market_id, offline_at_utc), fct_request (request_status, requested_at_local, market_id, dispatch_attempts) and dim_market (timezone). Deliverable: the three questions you ask before touching data, then a scoping note under 150 words naming the decision it serves, the one number that answers it with numerator, denominator and window, the first cut you run, and the date you return.

PracHub interview preparation framework
Is this an official Airbnb interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Airbnb. Rounds and questions reflect what candidates have reported, not a process Airbnb has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research
How difficult is the interview process, and how much preparation time should I plan for?

The interview loop is rigorous and thorough, reflecting the high standards of the organization. Most candidates dedicate between four to six weeks of focused preparation, particularly refreshing advanced SQL, experimental design nuances, and machine learning implementation.

PracHub interview research
What differentiates successful candidates from those who do not pass?

Successful candidates excel at bridging technical depth with product intuition. Rather than jumping straight into code or formulas, they take time to clarify ambiguity, state their assumptions, and tie their analytical approaches back to core business and user impact.

PracHub interview research
How are take-home assignments evaluated?

Take-home projects are evaluated on analytical rigor, code quality, and the clarity of your insights. Focus on producing clean, well-documented code and a concise summary that highlights your decision-making process rather than overwhelming the review team with excessive output.

PracHub interview research
What is the typical timeline from the initial recruiter screen to final offer?

The entire process generally spans three to four weeks from the initial recruiter conversation through the take-home challenge, technical screens, and virtual on-site rounds, though timelines can vary based on team scheduling and headcount needs.

PracHub interview research
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.