Via Transportation · Data Scientist
Updated · 2026-09-24

Via Transportation Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

A Data Scientist at Via Transportation sits at the intersection of complex algorithmic development and real-world urban logistics. Your work is fundamental to the company’s core mission: optimizing transit networks to make shared transportation more efficient, accessible, and sustainable. You aren't just building models; you are solving massive, dynamic optimization problems that directly impact how people move through cities.

How much statistics you need depends on the flavour of the seat. Experiment-facing work wants you deep enough to notice that repeated looks at accumulating data inflate the false positive rate of a fixed-sample test; modelling-facing work wants estimation and honest uncertainty intervals.

Via Transportation candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.

Trace an on-time miss to one nodeEvaluate forecasts at the lag ordering actually usesJoin daily snapshots to shipment events safely

35 min read

Practice 15 Data Scientist prompts
15Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Data Scientist at Via Transportation sits at the intersection of complex algorithmic development and real-world urban logistics. Your work is fundamental to the company’s core mission: optimizing transit networks to make shared transportation more efficient, accessible, and sustainable. You aren't just building models; you are solving massive, dynamic optimization problems that directly impact how people move through cities.

The role involves high-stakes technical challenges, such as demand prediction, fleet routing, and transit network design. You will collaborate closely with engineering, product, and operations teams to translate business needs into scalable data products. Because Via Transportation operates at a massive scale, your contributions—whether in machine learning, statistical modeling, or simulation—have immediate, tangible effects on the user experience and the company’s bottom line.

01

Initial Screening Call

reported

A screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.

What to demonstrate

  • Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
  • Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
  • Whether timeline, location and compensation expectations make the rest of the loop worth scheduling

How to prepare

  • Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
  • Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
  • Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
PracHub interview research ↗
02

Technical Interviews

reported

This round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.

What to demonstrate

  • Whether your row counts survive each join, and whether you notice on your own when they do not
  • Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
  • Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
  • Reaching a defensible answer inside the window instead of a refined one after it

How to prepare

  • Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
  • Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
  • Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
PracHub interview research ↗
03

Technical Take-Home Assignment

reported

Your submission is read asynchronously by someone who cannot ask you a clarifying question, and who will usually skim it before reading it properly. That changes what good looks like. The conclusion belongs near the top with the supporting analysis beneath it, and the code should run start to finish on a clean machine without a manual step you forgot to document. A reviewer forced to reconstruct your reasoning from the order of notebook cells is already discounting the work. The submissions that land are the ones where a busy reader gets the answer immediately and can verify it if they want to.

What to demonstrate

  • Whether the answer arrives before the methodology, so a reader who stops after the first page still has the recommendation
  • Whether the code runs end to end from the submitted files, with dependencies and data paths declared rather than assumed
  • Whether each chart is legible on its own, carrying axis labels and units, and exists to support a claim made in the text

How to prepare

  • Restructure a past analysis so the opening paragraph holds the recommendation and the number behind it, then check that nothing later in the document quietly contradicts it
  • Copy your own submission into an empty directory, run it in a clean environment, and fix everything that breaks; hidden local state is caught here or by the reviewer
  • For every chart, write the one sentence it is meant to prove, and delete the chart if you cannot write that sentence
PracHub interview research ↗
04

Deep-Dive Interviews

reported

Rounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.

What to demonstrate

  • Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
  • The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
  • Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options

How to prepare

  • Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
  • Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
  • Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
PracHub interview research ↗
05

On-Site Round

reported

Where a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.

What to demonstrate

  • Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
  • Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
  • Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
  • Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered

How to prepare

  • Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
  • Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
  • Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Treating shipped units as demand

Shipments are censored at available inventory: on a stockout day the recorded quantity is a supply ceiling, not customer intent, and orders that were never placed because the item showed as unavailable leave no row at all. A forecast fitted on that history learns the constraint, under-forecasts the fast movers that stock out most often, and produces the replenishment that causes the next stockout, so the error compounds in one direction rather than averaging out. The fix is to model demand_qty rather than shipped_qty, flag stockout days with stockout_flag and treat them as censored (fit with a censored likelihood, or estimate unconstrained demand from uncensored periods and comparable locations), and to report how much of the history was censored alongside any accuracy number.

02

Computing average inventory from period-end snapshots

Shipments cluster before period close, so the month-end on-hand position is systematically the lowest point of the month; turns computed against it are biased high and days of supply biased low, frequently by ten to twenty percent, and the bias grows precisely when close-period push is strongest. The same shape of error appears when a daily snapshot is joined to shipment events on date equality: a SKU with several legs on one day fans the snapshot out and multiplies the valued inventory. Average across every daily snapshot in the window for the denominator, and join snapshots to events with an explicit as-of condition and a row-count check at each grain before aggregating.

03

Accepting a metric definition without asking about the denominator

Pin down the denominator, the eligibility filter and the time window before computing anything: conversion rate per session, per user, per eligible user and per new user are four different numbers with different behaviour. Restate the definition in one sentence and get agreement before you analyse.

04

Reporting a mean for a heavy-tailed metric without saying what it hides

For spend, session length or items per order, a small fraction of units carries most of the total, so the mean has a wide standard error and one account can move it. Fix the handling before you see the result: cap or winsorise at a pre-declared percentile, and report the median or the share above a threshold next to the mean. Capping changes the estimand, so say which question the capped number answers, and check how much of any difference comes from the top 0.1 percent of units.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

12 technical prompts3 include a worked solution

Rides in a specific neighborhood are receiving lower-than-average rati…

medium
machine learning and modelling

Rides in a specific neighborhood are receiving lower-than-average ratings; how would you investigate and model this?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Check what information would not exist at prediction time, and exclude it.
Follow-up
  • Where could label leakage enter this setup?
  • What would you monitor after launch to know the model is still valid?

How would you approach a clustering problem involving massive NYC tran…

medium
machine learning and modelling

How would you approach a clustering problem involving massive NYC transit datasets?

Approach
  1. Set a baseline first, so any model has something honest to beat.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

Simulate service levels under a variable inbound lead time

mediumWorked solution
safety stockmonte carloservice level

Daily demand for one SKU at one node is Poisson with mean 40. Replenishment lead time is 7 days with probability 0.7, 10 days with 0.2, and 14 days with 0.1, drawn independently per order. The policy is continuous review: order Q = 400 whenever the inventory position drops to the reorder point. Size safety stock two ways for a nominal 95 percent target, once as z times sigma_D times sqrt(mean lead time) and once with the variable-lead-time term. Simulate both and report achieved cycle service level and achieved unit fill rate.

Approach
  1. Compute both safety stock numbers analytically before simulating, so the simulation is a check on arithmetic you already understand rather than the source of the answer.
  2. Track inventory position (on hand plus on order) for the reorder trigger and on hand separately for the stockout test. Triggering off on hand alone re-orders while stock is already in transit and silently changes the policy you are measuring.
  3. Define the two outcomes precisely and separately: cycle service level is the fraction of replenishment cycles containing at least one unit of unmet demand, while fill rate is total units shipped from stock over total units demanded. They answer different questions and z targets only the first.
  4. Before running anything, bound the achievable service level by hand from the lead-time mixture. Each lead-time branch is a separate Poisson demand, so cycle service level is the probability-weighted sum of P(Poisson(40 x L) <= reorder point) across the three branches, and a reorder point that cannot cover the 14-day branch caps the policy at 0.9 no matter how the simulation is tuned.
  5. Run enough cycles that Monte Carlo error is small against the gap you are trying to see. The standard error of a service level is sqrt(p*(1-p)/n_cycles), about 0.005 at p = 0.7 with 10,000 cycles, so the 20-point gap between the two policies is unambiguous and a 1-point one is not.
  6. Let unmet demand be lost rather than backordered, state that choice, and note that switching to full backordering raises measured fill rate for the same policy.
Worked solution 35 min
  1. Compute mean lead time 8.3 days, variance of lead time 5.01, sigma_D = sqrt(40) = 6.32, and mean demand 40 per day.
  2. Naive safety stock: 1.645 * 6.32 * sqrt(8.3) = 30 units, giving a reorder point of 332 + 30 = 362. Correct safety stock: 1.645 * sqrt(8.3 * 40 + 40^2 * 5.01) = 1.645 * 91.4 = 150 units, giving a reorder point of 482.
  3. Write one simulation loop parameterised by safety stock: draw daily Poisson demand, decrement on hand, record shortfall, trigger an order of 400 when position hits the reorder point, and schedule the receipt after a sampled lead time.
  4. Run 10,000 cycles per policy with a burn-in of at least two lead times discarded so the starting position does not flatter the first cycles.
  5. Report achieved cycle service level, achieved fill rate and the average on-hand position for each of the two safety stock settings.
EXPECTED RESULTNaive safety stock is 30 units and the variable-lead-time figure is 150 units, a factor of five. At Q = 400 with lost sales, the naive policy achieves a cycle service level of about 0.70 and the correct policy about 0.90, so the naive policy misses the nominal 0.95 by roughly 25 points and the correct one by 5. The naive ceiling is arithmetic, not simulation noise: with a reorder point of 362 the 7-day branch is essentially always covered, the 10-day branch survives with probability P(Poisson(400) <= 362) = 0.029, and the 14-day branch never does, so cycle service level cannot exceed 0.7 + 0.2*0.029 + 0.1*0 = 0.706, and reorder-point undershoot of about 20 units under daily demand arrivals pulls it to roughly 0.70. The correct policy tops out near 0.90 for the same structural reason: a reorder point of 482 covers the 10-day branch almost completely (P = 0.99997) but not the 14-day one (P = 0.0004), so that 0.1 of probability mass is near-pure loss. Fill rate is higher than cycle service level under both, about 0.92 against 0.70 and about 0.98 against 0.90, because expected units short per cycle (roughly 33 and 10) are small against Q = 400.
Follow-up
  • With Q = 100 against lead-time demand of 332 units, three or four orders are open at once. What happens to the meaning of a replenishment cycle, and to each of the two service measures?
  • Lead time and demand are correlated because the supplier is capacity constrained in peak weeks. What does that do to the formula you used?
  • How would you convert the extra 120 units of safety stock into an annual cost the business can weigh against the service gain?

Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Diagnostic, scored before you study anything
  • Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
  • Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
  • Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.

Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Largest gap: find the boundary rather than the subject
  • Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
  • Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
  • Re-solve one of them from memory four hours later, on paper, with nothing open.

Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.

Practice prompt ↗Practice prompt ↗
03Largest gap: drill the blocking sub-skill
  • Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
  • Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
  • Have someone else read your one-sentence rule and find the precondition you left out.

Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.

Practice prompt ↗Practice prompt ↗
04Second gap, plus maintenance on your strongest area
  • Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
  • Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
  • Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.

Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05The gap that is not a skill
  • Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
  • Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
  • Deliver one answer again with a hard 90-second limit to force structure before detail.

Deliverable: Two recordings with a counted improvement in time-to-first-question.

Practice prompt ↗Practice prompt ↗
06Retest under day-one conditions
  • Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
  • Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
  • Write which single block you would still lose the offer on.

Deliverable: A second scored rubric placed next to the first, with one named remaining risk.

Practice prompt ↗Practice prompt ↗
07Full loop under interview conditions
  • Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
  • Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
  • Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.

Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Interviewers here are not checking whether you can describe a project. They want the decision you made, why you made it under the information you had, and what changed afterwards that someone else could measure. A story that ends at 'I built a model' has no ending. Say what the model caused, or what you stopped doing because of it.

Describe your experience with SQL and data manipulation at scale.

medium
behavioural and stakeholder questions

Describe your experience with SQL and data manipulation at scale.

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Close with what you would do differently, concretely.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • How did you know the outcome was caused by your change?
  • What would you do differently if you ran that project again?

How do you handle missing or noisy data when dealing with real-time GP…

medium
behavioural and stakeholder questions

How do you handle missing or noisy data when dealing with real-time GPS or transit logs?

Approach
  1. Name the disagreement or constraint, and how you resolved it with evidence.
  2. Quantify the outcome, including what you would not claim credit for.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

Explain forecast uncertainty to a non-technical general manager

easy
communicating uncertaintyservice levelexecutive communication

A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.

Approach
  1. Convert the error band into the units the approver already thinks in: weeks of cover, pallets, or dollars at risk, not a percentage or a confidence interval.
  2. Answer the right-or-wrong question directly rather than deflecting: the forecast will be wrong, the size of the wrongness is what you measured, and the build is sized against that size.
  3. State precisely what the 97 percent buys: it is a cycle service level, the probability of not running out during one replenishment cycle, so roughly three cycles in a hundred see a stockout. It is not the share of units shipped from stock. Unit fill rate also depends on the replenishment quantity, and when that quantity is large relative to the standard deviation of lead-time demand, which is the normal case outside lot-for-lot ordering, the fill rate sits above the cycle service number, often above 99 percent at a 97 percent cycle target.
  4. Give the two-sided consequence in money: what the build costs to carry at the applicable cost of capital plus obsolescence risk on shelf-life items, against the margin at risk from the stockouts it prevents.
  5. Close with the decision you want and the trigger that would reverse it, for example a lag-7 bias check after four weeks that reopens the number.
Follow-up
  • The GM says just give me one number. What do you give, and what do you refuse to give?
  • How would your answer change if the items were frozen with a 90-day shelf life?
  • 01

    Describe your experience with SQL and data manipulation at scale.

  • 02

    How do you handle missing or noisy data when dealing with real-time GPS or transit logs?

  • 03

    A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.

PracHub interview preparation framework ↗
Is this an official Via Transportation interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Via Transportation. Rounds and questions reflect what candidates have reported, not a process Via Transportation has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
Is the take-home assignment really necessary?

Yes, it is a core component of the evaluation process at Via Transportation. It allows the team to assess your real-world coding ability, your approach to ambiguous problems, and your ability to communicate findings in a professional format.

PracHub interview research ↗
What is the best way to stand out during the interview?

Focus on your "business intuition." The most successful candidates are those who don't just build a model, but explain how that model solves a specific business problem and what the potential trade-offs are for the company.

PracHub interview research ↗
How long does the entire process take?

The process can vary, but it often spans several weeks due to the multiple interview rounds and the time allotted for the take-home challenge. It is best to maintain consistent communication with your recruiter throughout.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.