As a Data Scientist at Lyft, you sit at the epicenter of a dynamic, two-sided marketplace where millions of riders and drivers connect daily. Data science is not an auxiliary function at Lyft; it is the core driver of product strategy, automated decision-making, pricing mechanisms, and operational efficiency. Whether you join the Decisions track—focusing on product analytics, causal inference, and business strategy—or the Algorithms track—focusing on machine learning, optimization, and real-time decision systems—your work directly impacts key company metrics such as total rides completed, driver earnings, marketplace balance, and platform safety.
Data Scientists at Lyft work across high-impact business units including Rider & Safety, Mapping & Routing, Base Earnings, Growth, Lyft Ads, and Fulfillment. In these domains, you will solve complex, unstructured problems: evaluating how to match riders and drivers optimally, establishing dynamic pay policies, forecasting supply-demand imbalances, or mitigating platform safety risks. Because ride-hailing operates in real-time within complex physical environments, your analyses and models must account for spatial-temporal constraints, network spillover effects, and economic incentives.
The culture surrounding data science at Lyft is fast-paced, highly collaborative, and analytically rigorous. You will work side-by-side with product managers, software engineers, operations leads, and executive leadership to turn massive datasets into actionable strategic decisions or production-ready algorithmic features. To succeed, you must combine deep technical proficiency in statistics and SQL with exceptional business intuition, structured communication, and the ability to drive alignment across cross-functional teams.
Recruiter Conversation
reportedWhoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.
What to demonstrate
- Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
- Whether your reason for wanting the role points at the work itself rather than the company's reputation
- Whether your language signals the level being screened for: what you decided yourself versus what you were handed
How to prepare
- Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
- Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
- Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
Technical Phone Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Take-Home Data Challenge
reportedYour submission is read asynchronously by someone who cannot ask you a clarifying question, and who will usually skim it before reading it properly. That changes what good looks like. The conclusion belongs near the top with the supporting analysis beneath it, and the code should run start to finish on a clean machine without a manual step you forgot to document. A reviewer forced to reconstruct your reasoning from the order of notebook cells is already discounting the work. The submissions that land are the ones where a busy reader gets the answer immediately and can verify it if they want to.
What to demonstrate
- Whether the answer arrives before the methodology, so a reader who stops after the first page still has the recommendation
- Whether the code runs end to end from the submitted files, with dependencies and data paths declared rather than assumed
- Whether each chart is legible on its own, carrying axis labels and units, and exists to support a claim made in the text
How to prepare
- Restructure a past analysis so the opening paragraph holds the recommendation and the number behind it, then check that nothing later in the document quietly contradicts it
- Copy your own submission into an empty directory, run it in a clean environment, and fix everything that breaks; hidden local state is caught here or by the reviewer
- For every chart, write the one sentence it is meant to prove, and delete the chart if you cannot write that sentence
Final Round Onsite
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
8 candidate reports. Individual accounts describe a particular role and hiring cycle.
Lyft Senior Software Engineer Interview Experience — Messenger Design and a Next-Day Pass
The target level was Senior. The process was bizarre, but I'll leave that aside. Overall, the interviews weren't difficult. Surprisingly, the hardest coding problem was in the phone screen. Phone-screen coding: LeetCode 76. VO 1: The behavioral round. They asked about my past projects, which project I was proudest of, how I had moved projects forward with colleagues over the past two years, and s…
Read full experienceLyft Software Engineer Interview Experience: A transactional recruiter screen
The recruiter screen set a rough tone immediately. I came in for a backend engineering role, but the conversation felt overly transactional, as though boxes were being checked instead of discussing the technical impact I had delivered. I shared concrete examples from past work, including architecture decisions, optimizations, and engineering scenarios. I expected the conversation to connect those…
Read full experienceLyft Data Scientist interview: probability to causal modeling
I began with a recruiter screen of about 30 minutes. It was structured: they checked the basic boxes and explained what would happen next. The process then moved through a general data-science screen on probability, A/B testing, and business sense, followed by four deeper technical rounds. Those covered coding and algorithms, optimization, causal modeling, and business acumen tied back to product…
Read full experienceLyft Intern Software Engineer Interview Experience — Three Rounds, Rejected for No Headcount After Three Months
The OA had two parts. The first was responding to a PM's comments inside a doc, and the second was coding — more like they hand you a stripped-down existing project and have you edit the code according to review comments, not your typical LeetCode problem. VO1: one hour of coding, something like LRU. I basically finished within 30 minutes. The interviewer was really nice, and there weren't many f…
Read full experienceLyft Software Engineer Interview Experience — Onsite Loop with a Paginated Fetch Design and a Job Scheduler
Onsite. They gave me a long chunk of existing code and asked me to implement a function based on it. Essentially there's a fetch(page) function that returns the items on that page along with the next page's nextPage. We needed to implement fetch_n as a method on another class, so that we could pull n elements continuously. Each call continues extracting from wherever the last call left off, so th…
Read full experiencePracHub editorial advice for the preparation topics above.
Conditioning the analysis on completed orders
Wait-time distributions, price elasticities and rating models fit only on completed orders are conditioned on an outcome that the intervention itself changes. The requests that never matched, or that the consumer abandoned, are the population a liquidity fix targets, so excluding them biases every estimate toward the status quo and can flip the sign of a price elasticity. Any query starting FROM fct_order is already inside this trap; start from fct_request and left join.
Randomising individual consumers when supply is shared
A feature that makes treatment consumers book faster consumes the same idle providers the control consumers would have used, so the control group is degraded by the treatment and the measured lift overstates the market-level effect. The bias is largest precisely when supply is tight, which is when the feature is supposed to help, so the experiment is most misleading exactly where the decision matters. The fix is randomising the market or the time block (switchback) and clustering the variance at the randomisation unit, accepting far fewer effective units.
Reading a dozen metrics with no multiplicity control
Nominate one primary metric before launch and treat the rest as guardrails or exploratory, with Bonferroni or Benjamini-Hochberg applied when you intend to make claims from them. Twenty independent tests at 0.05 under the null produce at least one false positive about 64 percent of the time.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain how you would construct a confidence interval for the success …
Explain how you would construct a confidence interval for the success rate parameter $p$ in a binomial distribution, and how sample size impacts the interval width.
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Translate the result into the decision it informs, in one plain sentence.
- Say what the estimate is of, and over what population it generalises.
Follow-up
- How would you explain this result to someone who does not know statistics?
- What sample size would you need to detect an effect half this size?
Explain how to implement stratified k-fold cross-validation from scrat…
Explain how to implement stratified k-fold cross-validation from scratch in Python, and discuss why stratification is critical for imbalanced target variables.
Approach
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Net take and margin per market-month without ledger fan-out
You have orders (order_id, market_id, completed_at_utc, order_status, gross_booking_cents, provider_payout_cents, consumer_incentive_cents, provider_incentive_cents, tip_cents) and ledger (ledger_id, order_id, entry_type in charge, refund, chargeback, incentive, payout, adjustment, processing_fee; amount_cents signed positive into the platform; posted_at_utc; settlement_status). For completed orders, produce a market-month frame with net take, margin before support cost, and the order count. Attribute refunds, chargebacks and processing fees to the order's completion month, not the posting month, and drop ledger rows whose settlement_status is 'pending' or 'failed'.
Approach
- Aggregate the ledger to the order grain before touching orders: filter settlement_status, then pivot entry_type into one summed column each. Merging first and summing after multiplies gross_booking_cents by the number of ledger rows on that order.
- Merge the pivoted ledger onto orders with validate='m:1' so an unexpected duplicate raises instead of silently inflating every total.
- Net take = gross_booking - provider_payout - consumer_incentive - provider_incentive over completed orders only; tips are excluded on both sides because they pass through to the provider and never enter platform revenue.
- Add the signed ledger columns rather than subtracting absolute values: refunds, chargebacks, payouts and processing fees are negative under this convention, so margin = net_take + refund + chargeback + processing_fee. Verify the sign on one known order before trusting the aggregate.
- Group by market_id and completed_at_utc month, never posting month; that is the whole point of the attribution rule, and it is why a month's margin is not final until the chargeback window closes.
- Carry completed_orders in the output so per-order margin can be recomputed downstream without averaging an average.
Worked solution 30 min
- led = ledger[~ledger.settlement_status.isin(['pending','failed'])]; wide = led.pivot_table(index='order_id', columns='entry_type', values='amount_cents', aggfunc='sum', fill_value=0).
- comp = orders[orders.order_status == 'completed']; m = comp.merge(wide, on='order_id', how='left', validate='m:1').fillna({col: 0 for col in wide.columns}).
- net_take = gross_booking_cents - provider_payout_cents - consumer_incentive_cents - provider_incentive_cents.
- margin = net_take + refund + chargeback + processing_fee, using the signed ledger columns directly.
- Group by market_id and completed_at_utc.dt.to_period('M'), summing net_take and margin and counting order_id.
Follow-up
- A chargeback posts two months after completion and changes an already-reported month. How do you publish a metric that is not final, and what lag would you quote?
- Should the denominator be matched orders, completed orders, or completed-and-settled orders? Argue for one and name what it hides.
Given a log table of ride requests, write a query using SQL window fun…
Given a log table of ride requests, write a query using SQL window functions to calculate each driver's rolling 7-day acceptance rate and rank drivers by completion speed.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Write a SQL query to identify riders who took a ride to work in the mo…
Write a SQL query to identify riders who took a ride to work in the morning but did not book a return ride home with Lyft on the same day.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Fill rate by market-day without averaging hourly rates
fct_request holds one row per demand request: request_id, market_id, requested_at_utc, requested_at_local, matched_at_utc (NULL when the request never matched), terminal_at_utc, request_status. dim_market holds market_id, vertical, timezone. Produce SLA fill rate per market per local calendar day for the last 28 days. A request enters the numerator when matched_at_utc is not null and matched_at_utc minus requested_at_utc is at or under 120 seconds for vertical 'ride' and 600 seconds for 'delivery'. Exclude requests with request_status = 'abandoned_pre_match' whose terminal_at_utc is within 10 seconds of requested_at_utc. Return market_id, local_date, numerator, denominator, rate.
Approach
- Build the eligible population as a CTE first: cut the 28-day window, then drop the sub-10-second abandons. Applying the exclusion before any aggregation guarantees the numerator and the denominator are computed over the identical row set.
- Cut the window and take the bucket off the same clock. local_date is date(requested_at_local), and the window is local_date BETWEEN end_date - 27 AND end_date, where end_date is the last local calendar date that has closed in every market. Cutting on requested_at_utc instead moves the day boundary by up to 14 hours: the same query then returns 29 local dates for any market at a non-zero UTC offset, the first and last of them partial. A partial edge day holds a few hours of one market's overnight or evening traffic, so its rate is computed over a slice whose fill rate is nothing like the day's, and it is then read beside 27 whole days.
- Express the SLA as a CASE on dim_market.vertical rather than a hardcoded constant, then compute the numerator as COUNT(*) FILTER (WHERE matched_at_utc IS NOT NULL AND matched_at_utc - requested_at_utc <= sla).
- Aggregate once at market by local_date and divide summed numerator by summed denominator. If an hourly view is also wanted, compute the hourly counts and re-sum them to the day; do not average the hourly rate.
- Divide through NULLIF(denominator, 0). Under the CTE-first shape a market-day whose rows were all excluded produces no row at all rather than a zero denominator, so a gap in the dates means zero surviving requests, not a zero rate. Keep the NULLIF anyway: it is what stops a divide-by-zero the day someone moves the exclusion inside the aggregate instead of ahead of it.
Worked solution 20 min
- Count rows inside the 28 local-date window before and after the sub-10-second exclusion; the exclusion should remove a low single-digit percentage, not a third of the table.
- Write the CASE-based SLA and hand-check five matched requests against their computed time-to-match.
- Aggregate to market by local_date with COUNT() for the denominator and COUNT() FILTER (...) for the numerator, in one pass.
- Divide the summed counts with NULLIF on the denominator, then count distinct local_date per market before ordering the output by market_id, local_date.
Follow-up
- Fill rate is flat at the market-day level but down four points in the 17:00 to 19:00 local hours. Which cut do you run next, and what would confirm a supply cause rather than a demand spike?
- The 10-second exclusion removes six percent of requests in one market and half a percent in another. What do you check before trusting either market's rate?
Given a dataset of active riders and available drivers in a specific r…
Given a dataset of active riders and available drivers in a specific region, how would you design a matching framework to optimize driver utilization and rider wait times?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Walk through your approach to product metric design for the Lyft drive…
Walk through your approach to product metric design for the Lyft driver app home screen. What primary, secondary, and guardrail metrics would you track?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you evaluate the economic impact of lowering base fares vers…
How would you evaluate the economic impact of lowering base fares versus increasing targeted driver incentives during peak weekend hours?
Approach
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Drivers in a key market are converting from registration to their firs…
Drivers in a key market are converting from registration to their first ride at a lower rate than historical averages. How would you investigate this friction?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
You run an experiment on a new rider checkout UI and observe a statist…
You run an experiment on a new rider checkout UI and observe a statistically significant increase in conversion but a drop in total lifetime value (LTV). How do you decide whether to launch?
Approach
- State the primary metric and the minimum effect worth shipping, then size the test.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- What would you do if you could not randomise at all?
How would you design an A/B testing framework to evaluate a new driver…
How would you design an A/B testing framework to evaluate a new driver incentive structure without introducing market-level spillover bias between treatment and control groups?
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Decide the analysis before seeing data, including how long it runs and when you look.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
Estimate a single-market fee change without a control group
Regulation forces one market to cut its service fee on a known date. Randomisation is impossible and only that market is affected. You have 52 weeks of pre-period weekly series for every market: completed orders, active consumers, SLA fill rate, provider online hours and net take rate. Estimate the effect on completed orders per active consumer over the 12 weeks after the change. Specify the estimator, the donor pool and its exclusions, the pre-period fit criterion you would accept, and how you would produce a p-value with exactly one treated unit.
Approach
- Choose synthetic control as the estimator. Fit non-negative donor weights that sum to one, matching the treated market pre-period outcome path plus a small set of predictors such as vertical, launch age, price level and supply density. The convexity constraint is the point: it forbids extrapolation beyond the donor support, which an unrestricted regression fit would happily do.
- Build the donor pool by exclusion, not by convenience. Drop markets with their own policy or pricing change inside the window, and drop markets adjacent enough that providers or consumers can cross into the treated market within a session. A spilled-over donor is contaminated toward the treated path and shrinks the estimated effect toward zero.
- Set the fit criterion before looking at the post period. Require a small pre-period RMSPE relative to the effect size that would change the decision, and validate it out of sample by fitting weights on weeks 1 to 40 and checking fit on weeks 41 to 52, so the weights are not tuned to pre-period noise.
- Produce inference by permutation, because there is one treated unit and no cluster to compute a standard error over. Run the identical procedure treating each donor as if it were treated, compute the ratio of post-period to pre-period RMSPE for each, and rank the true treated market. With 30 donors the smallest attainable p-value is 1/31 = 0.032, so state that floor rather than reporting a conventional significance claim.
- Add a placebo in time: move the treatment date 12 weeks earlier and confirm the method produces a gap near zero. A method that manufactures an effect on a date when nothing happened cannot be trusted on the date when something did.
- Report the ratio decomposed into its parts. A fee cut moves completed orders and active consumers together, so completed orders per active consumer can stay flat while both components move substantially, and the decomposition is what the decision actually needs.
Worked solution 35 min
- Assemble the weekly panel and screen the donor pool for policy changes and geographic spillover, recording each exclusion and its reason.
- Fit donor weights on pre-period weeks 1 to 40 and validate the fit on weeks 41 to 52, reporting pre-period RMSPE.
- Compute the post-period gap between the treated market and its synthetic counterpart for each of the 12 weeks.
- Run placebo-in-space over every donor, rank the treated post/pre RMSPE ratio, and convert the rank into a permutation p-value with its 1/(N+1) floor.
- Run placebo-in-time at a fake treatment date 12 weeks early and confirm a near-zero gap.
- Decompose the ratio into completed orders and active consumers and report both alongside the ratio.
Follow-up
- Three markets get the same regulation on three different dates. What changes in the estimator, and what goes wrong with a naive two-way fixed-effects specification?
- Pre-period fit is excellent but the treated market is the largest in the pool. What should you suspect about the weights?
- How would you separate the effect of the fee change from the effect of the consumer-facing price change it caused?
Company take rate fell while every market's take rate rose
Monthly net take rate fell from 18.4% to 17.9% company-wide, yet all 46 markets individually improved or held. Net take rate is SUM(gross_booking_cents - provider_payout_cents - consumer_incentive_cents - provider_incentive_cents) / SUM(gross_booking_cents) over completed fct_order rows, per market per calendar month. dim_market carries launched_on, vertical and currency_code. Explain the arithmetic, quantify how much of the decline is mix and how much is within-market, and say which of the two numbers an operating team should be held to.
Approach
- Write the company metric as a weighted mean of market take rates, take_total = sum over m of w_m x take_m, where w_m is that market's share of gross bookings. The weights must be dollars, not order counts, because the metric is a ratio of dollars; weighting by orders yields a decomposition that does not reconcile.
- Apply the exact three-term shift-share: delta take = sum(delta w_m x take_m0) + sum(w_m0 x delta take_m) + sum(delta w_m x delta take_m). Report the interaction term rather than quietly folding it into one side.
- Rank markets by delta w_m x (take_m0 - take_total0) to name which cells are doing the mixing. Recently launched markets with heavy incentive intensity are the usual source, since their take rate is structurally lower while their share is growing fastest.
- Fix the currency basis before summing. fct_order amounts sit in the market's own currency per dim_market.currency_code, so any cross-market sum needs a rate; holding the rate at its month-zero value separates real weight shifts from FX movement.
- Publish both series: the raw company number, and a mix-adjusted number computed by holding w_m at a stated reference month. Say plainly which one operating teams are accountable for and which one the P&L actually sees.
Follow-up
- A market launched mid-month. Does it belong in either month's weights, and what does including it do to the interaction term?
- Leadership wants a single headline number. Which do you give them, and what do you put beside it?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.
Describe a time when you had to make a high-stakes business recommenda…
Describe a time when you had to make a high-stakes business recommendation using incomplete or ambiguous data. How did you structure your analysis?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Choose between three teams' requests with one analyst week
Three requests arrive the same morning and you have one analyst-week. Pricing wants a fee elasticity refresh for a change scheduled in ten days. Supply wants a churn model for approved providers who never opened a session. Finance wants the monthly contribution-margin restatement that attributes refunds and chargebacks in fct_money_movement to the order's completion month rather than the posting month. Deliverable: your ranking, the criterion behind it, the smallest useful version of each, and what you say to the two teams you rank below first.
Approach
- Rank on decision coupling rather than requester seniority or intrinsic interest: what decision hangs on this, on what date does it become useless, and how expensive is it to reverse if the answer is wrong.
- Notice the dependency before the ranking: the margin restatement changes the denominator of every unit-economics answer, including the elasticity work, so doing it second means redoing part of the first task.
- Unbundle each request into its smallest decision-bearing piece. The restatement is a change to the attribution date in one query, about half a day. The elasticity refresh needs a range and its preconditions, about two days. The churn model is the only item with no date and the longest build.
- Replace the churn model with the descriptive cut that may make it unnecessary: approved-provider time-to-first-session conversion by approval cohort and acquisition_channel, from dim_user provider_approved_at_utc against the first fct_supply_session.online_at_utc, one day. If conversion collapses in one channel or one market, the fix is operational and no model is needed.
- Communicate the ranking in writing where all three can see it, with the reason and the trigger that would reorder it, and give each deprioritised team a smaller concrete deliverable rather than a place in a queue.
Follow-up
- The supply lead escalates to your manager - what do you say, and what do you not say?
- What if the fee change's elasticity cannot be estimated observationally over the range they need?
- Which of the three would you drop entirely if you lost two days to an incident?
Disagree with a product lead about the randomisation unit
A product lead wants to ship a faster-dispatch change on the strength of a consumer-level A/B: 400,000 consumers, +6.1% completed orders per consumer, p below 0.001. You believe the design measures the wrong quantity, because treated consumers take the idle providers control consumers would otherwise have matched with. The lead's counter is that the sample is enormous and the p-value tiny. Deliverable: the argument you make including the direction of the bias, the evidence you can produce from existing data in two days, and the design you propose instead with its calendar cost.
Approach
- Concede the internal comparison and dispute the estimand: the test cleanly estimates a between-consumer contrast under a shared supply pool, which is not the market-level effect of shipping to everyone. Sample size does not touch this, because the bias does not shrink with n.
- State the direction and the mechanism: control consumers are degraded by treatment, so the contrast is inflated. The inflation is largest when idle supply is scarce, which is precisely the condition the feature is meant to help, so the test is most wrong where the decision matters most.
- Produce the fingerprint from data already in hand: split the measured lift by market-hour provider utilisation decile. Interference predicts the lift rises with utilisation; a genuine effect that does not steal supply does not have to. Pair it with control-arm fill rate against a pre-period baseline in high-utilisation hours, stating the precondition that this comparison is only informative if the pre-period is seasonally comparable or an untested market is available.
- Propose the switchback concretely: block length longer than a typical order duration so carryover does not leak across the boundary, a burn-in discarded after each switch, randomisation at the market-hour, and variance clustered at the block.
- Price the design honestly in calendar time so the lead can trade it off, and offer an interim: ship to one market with the rest held out, which is slower to read but not biased in the same direction.
Follow-up
- If the lift does not rise with utilisation, what does that tell you, and would you then ship?
- How do you choose block length when order durations have a long right tail?
- What variance reduction still works under interference, and what does it assume?
- 01
Describe a time when you had to make a high-stakes business recommendation using incomplete or ambiguous data. How did you structure your analysis?
- 02
Three requests arrive the same morning and you have one analyst-week. Pricing wants a fee elasticity refresh for a change scheduled in ten days. Supply wants a churn model for approved providers who never opened a session. Finance wants the monthly contribution-margin restatement that attributes refunds and chargebacks in fct_money_movement to the order's completion month rather than the posting month. Deliverable: your ranking, the criterion behind it, the smallest useful version of each, and what you say to the two teams you rank below first.
- 03
A product lead wants to ship a faster-dispatch change on the strength of a consumer-level A/B: 400,000 consumers, +6.1% completed orders per consumer, p below 0.001. You believe the design measures the wrong quantity, because treated consumers take the idle providers control consumers would otherwise have matched with. The lead's counter is that the sample is enormous and the p-value tiny. Deliverable: the argument you make including the direction of the bias, the evidence you can produce from existing data in two days, and the design you propose instead with its calendar cost.
Is this an official Lyft interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Lyft. Rounds and questions reflect what candidates have reported, not a process Lyft has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the take-home data challenge, and how much time should I expect to spend on it?
The take-home challenge is comprehensive and requires analyzing an open-ended dataset, conducting exploratory analysis, drawing strategic conclusions, and building presentation slides. While official guidelines estimate 5 to 8 hours, many successful candidates spend significant focus over 2–3 days refining their presentation clarity and methodology.
PracHub interview research ↗What differentiates successful candidates in the Lyft Data Science loop?
Successful candidates balance technical depth with clear, concise communication. They do not just deliver correct numbers or code; they clearly articulate the business implications, acknowledge modeling assumptions, and present structured trade-offs that resonate with cross-functional product leaders.
PracHub interview research ↗What is the work model and location policy for Data Scientists at Lyft?
Lyft operates primarily on a hybrid work schedule. Salaried team members in hybrid roles are expected to work in an designated office three days per week (typically Mondays, Wednesdays, and Thursdays), with flexibility to work fully remotely for up to four weeks per year.
PracHub interview research ↗How long does the entire interview process take from screen to offer?
The typical end-to-end timeline ranges between 3 to 5 weeks. However, because take-home assignments and final loop panel scheduling depend on candidate and team availability, candidate experience reports indicate that process duration can occasionally extend to 6 weeks.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22