A Data Scientist at Via Transportation sits at the intersection of complex algorithmic development and real-world urban logistics. Your work is fundamental to the company’s core mission: optimizing transit networks to make shared transportation more efficient, accessible, and sustainable. You aren't just building models; you are solving massive, dynamic optimization problems that directly impact how people move through cities.
The role involves high-stakes technical challenges, such as demand prediction, fleet routing, and transit network design. You will collaborate closely with engineering, product, and operations teams to translate business needs into scalable data products. Because Via Transportation operates at a massive scale, your contributions—whether in machine learning, statistical modeling, or simulation—have immediate, tangible effects on the user experience and the company’s bottom line.
Initial Screening Call
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Interviews
reportedThis round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.
What to demonstrate
- Whether your row counts survive each join, and whether you notice on your own when they do not
- Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
- Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
- Reaching a defensible answer inside the window instead of a refined one after it
How to prepare
- Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
- Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
- Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
Technical Take-Home Assignment
reportedYour submission is read asynchronously by someone who cannot ask you a clarifying question, and who will usually skim it before reading it properly. That changes what good looks like. The conclusion belongs near the top with the supporting analysis beneath it, and the code should run start to finish on a clean machine without a manual step you forgot to document. A reviewer forced to reconstruct your reasoning from the order of notebook cells is already discounting the work. The submissions that land are the ones where a busy reader gets the answer immediately and can verify it if they want to.
What to demonstrate
- Whether the answer arrives before the methodology, so a reader who stops after the first page still has the recommendation
- Whether the code runs end to end from the submitted files, with dependencies and data paths declared rather than assumed
- Whether each chart is legible on its own, carrying axis labels and units, and exists to support a claim made in the text
How to prepare
- Restructure a past analysis so the opening paragraph holds the recommendation and the number behind it, then check that nothing later in the document quietly contradicts it
- Copy your own submission into an empty directory, run it in a clean environment, and fix everything that breaks; hidden local state is caught here or by the reviewer
- For every chart, write the one sentence it is meant to prove, and delete the chart if you cannot write that sentence
Deep-Dive Interviews
reportedRounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.
What to demonstrate
- Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
- The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
- Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options
How to prepare
- Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
- Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
- Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
On-Site Round
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub editorial advice for the preparation topics above.
Treating shipped units as demand
Shipments are censored at available inventory: on a stockout day the recorded quantity is a supply ceiling, not customer intent, and orders that were never placed because the item showed as unavailable leave no row at all. A forecast fitted on that history learns the constraint, under-forecasts the fast movers that stock out most often, and produces the replenishment that causes the next stockout, so the error compounds in one direction rather than averaging out. The fix is to model demand_qty rather than shipped_qty, flag stockout days with stockout_flag and treat them as censored (fit with a censored likelihood, or estimate unconstrained demand from uncensored periods and comparable locations), and to report how much of the history was censored alongside any accuracy number.
Computing average inventory from period-end snapshots
Shipments cluster before period close, so the month-end on-hand position is systematically the lowest point of the month; turns computed against it are biased high and days of supply biased low, frequently by ten to twenty percent, and the bias grows precisely when close-period push is strongest. The same shape of error appears when a daily snapshot is joined to shipment events on date equality: a SKU with several legs on one day fans the snapshot out and multiplies the valued inventory. Average across every daily snapshot in the window for the denominator, and join snapshots to events with an explicit as-of condition and a row-count check at each grain before aggregating.
Accepting a metric definition without asking about the denominator
Pin down the denominator, the eligibility filter and the time window before computing anything: conversion rate per session, per user, per eligible user and per new user are four different numbers with different behaviour. Restate the definition in one sentence and get agreement before you analyse.
Reporting a mean for a heavy-tailed metric without saying what it hides
For spend, session length or items per order, a small fraction of units carries most of the total, so the mean has a wide standard error and one account can move it. Fix the handling before you see the result: cap or winsorise at a pre-declared percentile, and report the median or the share above a threshold next to the mean. Capping changes the estimand, so say which question the capped number answers, and check how much of any difference comes from the top 0.1 percent of units.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Rides in a specific neighborhood are receiving lower-than-average rati…
Rides in a specific neighborhood are receiving lower-than-average ratings; how would you investigate and model this?
Approach
- Say how the offline result would be validated online before it is trusted.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
How would you approach a clustering problem involving massive NYC tran…
How would you approach a clustering problem involving massive NYC transit datasets?
Approach
- Set a baseline first, so any model has something honest to beat.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Simulate service levels under a variable inbound lead time
Daily demand for one SKU at one node is Poisson with mean 40. Replenishment lead time is 7 days with probability 0.7, 10 days with 0.2, and 14 days with 0.1, drawn independently per order. The policy is continuous review: order Q = 400 whenever the inventory position drops to the reorder point. Size safety stock two ways for a nominal 95 percent target, once as z times sigma_D times sqrt(mean lead time) and once with the variable-lead-time term. Simulate both and report achieved cycle service level and achieved unit fill rate.
Approach
- Compute both safety stock numbers analytically before simulating, so the simulation is a check on arithmetic you already understand rather than the source of the answer.
- Track inventory position (on hand plus on order) for the reorder trigger and on hand separately for the stockout test. Triggering off on hand alone re-orders while stock is already in transit and silently changes the policy you are measuring.
- Define the two outcomes precisely and separately: cycle service level is the fraction of replenishment cycles containing at least one unit of unmet demand, while fill rate is total units shipped from stock over total units demanded. They answer different questions and z targets only the first.
- Before running anything, bound the achievable service level by hand from the lead-time mixture. Each lead-time branch is a separate Poisson demand, so cycle service level is the probability-weighted sum of P(Poisson(40 x L) <= reorder point) across the three branches, and a reorder point that cannot cover the 14-day branch caps the policy at 0.9 no matter how the simulation is tuned.
- Run enough cycles that Monte Carlo error is small against the gap you are trying to see. The standard error of a service level is sqrt(p*(1-p)/n_cycles), about 0.005 at p = 0.7 with 10,000 cycles, so the 20-point gap between the two policies is unambiguous and a 1-point one is not.
- Let unmet demand be lost rather than backordered, state that choice, and note that switching to full backordering raises measured fill rate for the same policy.
Worked solution 35 min
- Compute mean lead time 8.3 days, variance of lead time 5.01, sigma_D = sqrt(40) = 6.32, and mean demand 40 per day.
- Naive safety stock: 1.645 * 6.32 * sqrt(8.3) = 30 units, giving a reorder point of 332 + 30 = 362. Correct safety stock: 1.645 * sqrt(8.3 * 40 + 40^2 * 5.01) = 1.645 * 91.4 = 150 units, giving a reorder point of 482.
- Write one simulation loop parameterised by safety stock: draw daily Poisson demand, decrement on hand, record shortfall, trigger an order of 400 when position hits the reorder point, and schedule the receipt after a sampled lead time.
- Run 10,000 cycles per policy with a burn-in of at least two lead times discarded so the starting position does not flatter the first cycles.
- Report achieved cycle service level, achieved fill rate and the average on-hand position for each of the two safety stock settings.
Follow-up
- With Q = 100 against lead-time demand of 332 units, three or four orders are open at once. What happens to the meaning of a replenishment cycle, and to each of the two service measures?
- Lead time and demand are correlated because the supplier is capacity constrained in peak weeks. What does that do to the formula you used?
- How would you convert the extra 120 units of safety stock into an annual cost the business can weigh against the service gain?
Weekly unit fill rate by ship-from node
fct_order_line holds one row per customer order line per ship-from node in its terminal state, with ordered_qty, shipped_qty, requested_ship_date, actual_ship_at_utc, line_status and ship_from_location_id; dim_location carries location_code. Produce weekly first-pass unit fill rate by node. Numerator is shipped_qty on lines shipped on or before requested_ship_date; denominator is ordered_qty on every line requested that week except line_status = 'cancelled_by_customer'. Return node, week, numerator, denominator and rate, plus one network total row. State in the output how substituted lines are counted.
Approach
- Anchor the query at order-line grain filtered on requested_ship_date, never at shipments: a line that was never filled has no shipment row, and starting from shipments deletes exactly the failures the metric exists to count.
- Build the numerator as a conditional SUM over the same row set, SUM(CASE WHEN actual_ship_at_utc IS NOT NULL AND its ship date <= requested_ship_date THEN shipped_qty ELSE 0 END), so short and backordered lines stay in the denominator instead of vanishing behind a WHERE clause.
- Exclude only line_status = 'cancelled_by_customer'. A line cancelled for lack of supply is a service failure and belongs in the denominator at full ordered_qty.
- Group by node and by DATE_TRUNC('week', requested_ship_date), then build the network row by re-summing numerator and denominator across nodes, not by averaging the node rates.
- Declare the substitution rule explicitly in one column or one comment: lines with substituted_sku_id NOT NULL count toward the numerator only if substitution is an accepted fill in the service definition.
Worked solution 20 min
- Filter fct_order_line to requested_ship_date inside the window and line_status <> 'cancelled_by_customer'; count the rows removed by that filter and keep the count.
- Compute num = SUM(CASE WHEN actual_ship_at_utc::date <= requested_ship_date THEN shipped_qty ELSE 0 END) and den = SUM(ordered_qty) in the same SELECT.
- GROUP BY ship_from_location_id, DATE_TRUNC('week', requested_ship_date) and join dim_location for location_code.
- Add the network row from a second aggregate over the same base set: SUM(num) / SUM(den).
- Report the excluded-line count and the substitution rule beside the rate so the number is auditable.
Follow-up
- actual_ship_at_utc is UTC but requested_ship_date is a local calendar date at the node. How does the comparison change for a node at UTC+9, and which direction does the error run?
- Fill rate rose two points while a node's short_reason_code mix moved from 'no_stock' toward 'credit_hold'. Is that a supply improvement?
- How would you hold SKU mix fixed so the network number is comparable month over month?
Stockout streaks of three days or more per cell
fct_inventory_daily carries one row per sku_id x location_id x inventory_date for every active pair, including days with nothing on hand, plus stockout_flag marking days where available_qty reached zero. Return every run of consecutive stockout days lasting three days or more: sku, location, start date, end date, run length, and demand_qty summed across the run. Return alongside it, per sku-location, the count of such runs in the window. Verify that the daily row series is contiguous for each pair before relying on date arithmetic, and say what you do about pairs that fail.
Approach
- Filter to stockout_flag = TRUE, then assign ROW_NUMBER() OVER (PARTITION BY sku_id, location_id ORDER BY inventory_date).
- Form the island key as inventory_date - rn * INTERVAL '1 day'. Two rows share a key only when their dates are exactly consecutive, so the key cannot weld separate runs together; the exposure runs the other way. A day with no snapshot row inside a true stockout breaks that run into fragments, and any fragment under three days is deleted by the threshold, so the run leaves the output rather than appearing with a wrong length.
- Prove contiguity before trusting the key: per sku-location compare COUNT(DISTINCT inventory_date) against (MAX(inventory_date) - MIN(inventory_date) + 1), and separately confirm no pair carries two rows on one date, since a duplicate date advances rn without advancing the calendar and shifts every key after it. Pairs that fail get a generated date spine: LEFT JOIN generate_series over the window so each day is labelled stockout, in stock, or no snapshot, then apply a stated rule for the no-snapshot days instead of letting the key make that call silently. The conservative rule is to break the run at the hole and mark both fragments gap_adjacent.
- Group by sku, location and island key for MIN and MAX date, COUNT(*) as length and SUM(demand_qty) as the demand exposed to the stockout, keeping runs of length >= 3.
- Aggregate once more per sku-location for the run count, and mark runs touching either window edge as censored, since their true length is unknown from this window alone.
Follow-up
- The longest runs cluster on X-class A items. What does that combination suggest to look at before touching the safety stock policy?
- How would you estimate lost demand across a run, given that demand_qty records the orders that were placed but not the orders nobody bothered to place?
- The snapshot is taken at each location's local end of day and the pairs span several timezones. What breaks, and what stays true?
Given a dataset of bus routes, how would you determine the most effici…
Given a dataset of bus routes, how would you determine the most efficient pathing?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you design a metric to evaluate whether to expand or contrac…
How would you design a metric to evaluate whether to expand or contract a for-hire-vehicle service in a new urban area?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
If you notice a decline in driver acquisition, what data would you gat…
If you notice a decline in driver acquisition, what data would you gather and how would you structure your analysis?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you define "efficiency" for a ride-sharing platform, and how…
How would you define "efficiency" for a ride-sharing platform, and how would you aggregate that data to support business decisions?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Walk me through the technical details and statistical strategy used in…
Walk me through the technical details and statistical strategy used in your past projects.
Approach
- Clarify what is being asked and what a complete answer would contain.
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
A planner scorecard that aggregation cannot flatter
Demand planners will be scored on forecast accuracy. fct_inventory_daily carries demand_qty, forecast_qty_lag7, forecast_qty_lag28 and stockout_flag at sku by location by date, with NULL forecasts before a series existed. Specify the scorecard: error metric, grain, lag, exclusion policy and guardrail. Then show the three ways a planner raises the headline without forecasting better, and state which single number you would prefer to accuracy if you could publish only one.
Approach
- Fix the grain and lag to the decision. The ordering policy consumes a forecast at one horizon for one sku-location-week cell, so the score is computed there; a number quoted at any other grain is not the number the policy used and cannot be compared to it.
- Choose WMAPE, sum of absolute errors over sum of actuals, or a pinball loss at the service quantile the policy targets. Rule out MAPE explicitly: it is undefined at zero actuals, which is most days for a C item at a forward-stocking location, and it is asymmetric where it is defined, so optimising it drives systematic under-forecasting.
- Show why aggregation flatters, and show it as arithmetic rather than opinion. The absolute value of a sum is at most the sum of absolute values, so errors cancelling across SKUs, locations or weeks make aggregated WMAPE weakly lower than the disaggregated figure on the same denominator, with no forecasting improvement involved.
- Name the other two routes: scoring at a shorter lag than ordering actually uses, and exclusion drift, where NULL-forecast rows, zero-demand cells or stockout days are dropped without publishing the excluded count, quietly removing the hardest cells from the sample.
- Set the guardrail on signed bias, weighted mean percentage error over the same window and cells, reported per ABC/XYZ cell. A series can be inaccurate but unbiased, or accurate on average and persistently short, and only the second wrecks inventory in one direction.
- Answer the last question directly: prefer the outcome that accuracy is a means to, achieved fill rate at a given inventory at cost per cell, because that is what the business buys, and accuracy that does not move it is not worth paying for.
Worked solution 40 min
- Build the scored panel at sku by location by week, summing demand_qty and taking the forecast at the lag the ordering policy uses, excluding NULL-forecast cells and counting every exclusion.
- Compute WMAPE and signed WMPE per planner per ABC/XYZ cell by re-summing numerator and denominator, never by averaging cell ratios.
- Recompute the same WMAPE at region and at month to demonstrate the mechanical improvement, and record the difference as the aggregation premium.
- Recompute with and without stockout days and with and without zero-demand cells, recording each shift as the exclusion premium.
- Publish WMAPE, WMPE, the excluded-row count and the paired inventory-and-service outcome, with grain and lag printed on the page rather than held in a footnote.
Follow-up
- Two planners have identical WMAPE and opposite bias. Which one is doing more damage, and to what?
- The scorecard is per planner but forecasts are hierarchical and reconciled. Whose number is the reconciled one, and who gets charged for the reconciliation loss?
Plant schedule instability generated two echelons downstream
A plant reports that the replenishment orders it receives swing far more than end-customer demand, and asks for a larger finished-goods buffer. You have fct_order_line for customer demand at store and forward-stock nodes, fct_inventory_daily (on_order_qty, demand_qty per sku-location-date), dim_location (echelon, parent_location_id) and dim_sku (units_per_case, cases_per_pallet). Measure amplification at each echelon, identify the node whose ordering rule generates it, and price the buffer request against fixing that rule. Deliverable: a per-echelon amplification table and a recommendation.
Approach
- Build two series per node at a fixed weekly bucket per SKU: the quantity the node ordered from its parent_location_id, and the demand it received from its children or from customers. Both series must cover the same SKU, the same weeks and the same node set, or the ratio between them means nothing.
- Compute amplification per node as Var(orders placed) / Var(demand received), and report the squared coefficient-of-variation form alongside it. Mean volumes differ by an order of magnitude between echelons, so a raw variance ratio is not comparable across nodes of different size while the CV-squared ratio is scale-free.
- Read the table echelon by echelon and find where the ratio steps up. The generating node is where it first exceeds one by a wide margin, not the node that feels the pain: amplification observed at the plant is the product of every ratio below it along that chain.
- Name the mechanism at that node rather than asserting bullwhip generically. Minimum order quantity and lot-size multiple, the review period, batching up to a full truckload, and forecast-driven ordering that reacts to its own safety stock recalculation each leave a distinct signature in the order-size histogram.
- Price both options rather than choosing one. The buffer request converts to units, then to cost at standard_cost_cents and to carrying cost at the applicable cost of capital. The rule change converts to more frequent, smaller replenishment and therefore more freight per unit. Present the trade with both numbers attached.
Follow-up
- Orders cluster on exact truckload multiples. Would you break the batching rule, and what does that cost in freight per unit?
- How would you test an ordering-rule change when there are twelve nodes, they draw on the same stock and the same outbound loads, and one replenishment cycle is three weeks?
- Part of the plant's variance is its own production batching. How do you separate inbound amplification from what the plant creates for itself?
Instead of guessing where the week should go, day one measures it under a fixed rubric and allocates the remaining hours in proportion to the gaps. The method is deliberately rigid: the allocation is written down before any studying starts and is not renegotiated when a topic turns out to be unpleasant.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 100-minute timed diagnostic in four blocks: 30 minutes of SQL across three prompts, 25 minutes of short-answer statistics, 25 minutes on one modelling or case prompt, and 20 minutes delivering one behavioural story aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct, and 0 is stuck, grading the output rather than how the attempt felt.
- Allocate the hours for days two to five roughly in proportion to 3 minus the score in each block, write the allocation down, and commit to not revising it midweek.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Break the weakest area into five named sub-skills (for query work: grain control, window frames, date arithmetic, set logic with NULLs, and reading a query plan) and rate each one, so the rest of the week targets a sub-skill instead of a subject.
- Solve three problems chosen to sit just above where the rating drops off, and for each write the first move you failed to make.
- Re-solve one of them from memory four hours later, on paper, with nothing open.
Deliverable: A five-item sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗03Largest gap: drill the blocking sub-skill
- Do eight short repetitions of the same shape rather than eight different problems, so what you practise is the pattern and not the puzzle.
- Write the rule you now hold in one sentence, then test it against a case built to break it: a ranking function over a column with ties, or a two-sample test on observations that are obviously dependent.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on your strongest area
- Run the same sub-skill map and boundary protocol on the second-largest gap, compressed into half the day.
- Spend 25 timed minutes on your strongest area to stop it decaying, choosing the hardest problem you can still finish rather than an easy warm-up.
- Compare how the two areas fail: whether you lose time on recall, on setup, or on arithmetic, because the fix differs for each.
Deliverable: A second sub-skill map plus a one-line diagnosis of how each area fails you.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record yourself answering one technical and one behavioural prompt, then count two things in the playback: how many seconds before your first clarifying question, and how many sentences you started without knowing where they ended.
- Rewrite your three most-used stock phrases into shorter versions, and practise saying "I do not know, here is how I would find out" without softening it into a guess.
- Deliver one answer again with a hard 90-second limit to force structure before detail.
Deliverable: Two recordings with a counted improvement in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 100-minute diagnostic structure with new prompts of comparable difficulty and score it on the identical rubric.
- Compare block by block, and for any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write which single block you would still lose the offer on.
Deliverable: A second scored rubric placed next to the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a 60-minute mock covering the two blocks that moved least, with an interviewer instructed to interrupt and change direction.
- Write your recovery script for the moment you go blank: restate the question, state your assumption, name the first thing you would check.
- Reduce the week to the rule statements you wrote, each with its preconditions attached, then say every one of them out loud without reading it and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive being interrupted.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Interviewers here are not checking whether you can describe a project. They want the decision you made, why you made it under the information you had, and what changed afterwards that someone else could measure. A story that ends at 'I built a model' has no ending. Say what the model caused, or what you stopped doing because of it.
Describe your experience with SQL and data manipulation at scale.
Describe your experience with SQL and data manipulation at scale.
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Close with what you would do differently, concretely.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
How do you handle missing or noisy data when dealing with real-time GP…
How do you handle missing or noisy data when dealing with real-time GPS or transit logs?
Approach
- Name the disagreement or constraint, and how you resolved it with evidence.
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Explain forecast uncertainty to a non-technical general manager
A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.
Approach
- Convert the error band into the units the approver already thinks in: weeks of cover, pallets, or dollars at risk, not a percentage or a confidence interval.
- Answer the right-or-wrong question directly rather than deflecting: the forecast will be wrong, the size of the wrongness is what you measured, and the build is sized against that size.
- State precisely what the 97 percent buys: it is a cycle service level, the probability of not running out during one replenishment cycle, so roughly three cycles in a hundred see a stockout. It is not the share of units shipped from stock. Unit fill rate also depends on the replenishment quantity, and when that quantity is large relative to the standard deviation of lead-time demand, which is the normal case outside lot-for-lot ordering, the fill rate sits above the cycle service number, often above 99 percent at a 97 percent cycle target.
- Give the two-sided consequence in money: what the build costs to carry at the applicable cost of capital plus obsolescence risk on shelf-life items, against the margin at risk from the stockouts it prevents.
- Close with the decision you want and the trigger that would reverse it, for example a lag-7 bias check after four weeks that reopens the number.
Follow-up
- The GM says just give me one number. What do you give, and what do you refuse to give?
- How would your answer change if the items were frozen with a 90-day shelf life?
- 01
Describe your experience with SQL and data manipulation at scale.
- 02
How do you handle missing or noisy data when dealing with real-time GPS or transit logs?
- 03
A general manager who does not use statistics must approve a 2.1 million dollar finished-goods build. Your recommendation rests on a demand forecast with a wide error band on the Z-class items and on a 97 percent cycle service target. You get two minutes, no formulas, and the question you will get back is whether the forecast is right. Deliver the explanation as you would say it aloud, including how you describe the error band in physical terms and what the 97 percent number does and does not promise.
Is this an official Via Transportation interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Via Transportation. Rounds and questions reflect what candidates have reported, not a process Via Transportation has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗Is the take-home assignment really necessary?
Yes, it is a core component of the evaluation process at Via Transportation. It allows the team to assess your real-world coding ability, your approach to ambiguous problems, and your ability to communicate findings in a professional format.
PracHub interview research ↗What is the best way to stand out during the interview?
Focus on your "business intuition." The most successful candidates are those who don't just build a model, but explain how that model solves a specific business problem and what the potential trade-offs are for the company.
PracHub interview research ↗How long does the entire process take?
The process can vary, but it often spans several weeks due to the multiple interview rounds and the time allotted for the take-home challenge. It is best to maintain consistent communication with your recruiter throughout.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22