A Data Scientist at May Mobility plays a pivotal role in shaping the future of autonomous vehicle (AV) technology. Operating at the intersection of cutting-edge robotics, machine learning, and transportation logistics, you will transform massive streams of raw vehicle telemetry, sensor logs, and operational data into actionable insights. Your work directly influences how autonomous shuttles navigate complex urban environments, optimize their routes, and ensure passenger safety.
The impact of this role is profound. Because May Mobility is in an active research and development (R&D) phase, you will not simply maintain existing pipelines; you will design foundational metrics and analytical frameworks from scratch. Whether you are analyzing rider demand patterns, optimizing fleet deployment, or defining safety-critical performance indicators, your models and analyses will drive strategic decisions across engineering, product, and operations teams.
This position is ideal for those who thrive in high-ambiguity environments. You will work with complex, unstructured datasets that represent real-world physical interactions. If you are passionate about applying statistical rigor to physical-world problems and want to see your algorithms directly impact autonomous fleets on public roads, the role offers an unparalleled opportunity for technical ownership and visible real-world impact.
Recruiter Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
Technical Take-Home Assessment
reportedA take-home is graded as an argument, not as a notebook. Somebody reads the submission without you in the room, so every choice has to survive on the page: why the question was framed this way, and what was deliberately left out. The gap between a strong and a weak submission is almost never model quality. It is whether the writeup names the specific question it answers and commits to a recommendation, including what evidence would overturn it. A high-accuracy model attached to no conclusion reads as effort that stopped before the decision.
What to demonstrate
- Whether the question you answered is stated outright, and whether it is the question the prompt posed rather than an easier neighbour of it
- Whether the recommendation is specific enough to act on, with the uncertainty attached to it instead of parked in a caveats section at the end
- Whether analytical choices such as the metric definition, the population filter and the time window are justified in the prose, not merely visible in code
How to prepare
- Take a dataset you have already worked with, write the one-paragraph conclusion first, then check whether the analysis you were planning actually supports it and cut whatever does not
- Practise stating a metric in one sentence that fixes the population, the time window and the denominator, then confirm your query computes exactly that sentence and nothing adjacent to it
- Hand a draft to someone outside the problem and ask them to tell you back what you recommended and why; anything they cannot recover is not on the page yet
Interviews with Hiring Manager
reportedMuch of this round runs on your own history, but the manager is not collecting a project list. They are working out what it is like when something goes wrong on your watch: how late the bad news tends to arrive, and whether a number you hand over has been checked by anyone including you. That is why the strongest material is a project where you can describe the part that did not work and what it cost. A result you cannot take full responsibility for, however clean, gives them nothing to trust you with afterwards.
What to demonstrate
- Whether you volunteer the limits of a result you are proud of, or wait to be pushed onto them
- How errors surfaced in your past work, and whether you or somebody else found them
- Whether the scope you claim matches the level of detail you can still produce about it
- What you did the first time a stakeholder acted on something of yours that turned out to be wrong
How to prepare
- Rebuild one headline figure from memory down to the join and the filter, so a question about the denominator does not stall the conversation
- For each project you raise, write the sentence you would say to someone who had already acted on a number that later turned out wrong
- Mark which parts of a project were yours and which belonged to other people, and state that boundary yourself before anyone asks
Interviews with Data Science Team
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
PracHub editorial advice for the preparation topics above.
Averaging rates across SKU-locations instead of re-summing
Fill rate, turns, OEE and on-time rate are all ratios whose denominators differ by orders of magnitude between cells, so an unweighted mean gives a slow-moving C item at a small node the same vote as a high-volume A item at a national node. The blended figure then moves whenever the portfolio mix moves, and it can improve in every cell while the company-level ratio worsens, or the reverse, which is Simpson's paradox with a warehouse attached. Always sum numerator and denominator to the reporting level and divide once, and when a rate must be compared across nodes, standardise on a fixed SKU mix before reading anything into the difference.
Sizing safety stock as z times sigma_D times the square root of lead time
That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.
Generalising beyond the population the sample actually supports
State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.
SQL that silently fans out on a one-to-many join
State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How would you define and measure "ride comfort" using vehicle accelero…
How would you define and measure "ride comfort" using vehicle accelerometer and gyroscope data?
Approach
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Write down the assumption the method needs before you use the method.
Follow-up
- How would you explain this result to someone who does not know statistics?
- Which assumption here is most likely to be violated in practice?
How do you determine if a sample size is statistically significant whe…
How do you determine if a sample size is statistically significant when testing autonomous shuttle performance in a new, low-traffic market?
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Say what the estimate is of, and over what population it generalises.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
If we introduce a new routing algorithm in a specific city, how would …
If we introduce a new routing algorithm in a specific city, how would you design an A/B test to measure its impact on passenger wait times?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
Measure demand amplification hop by hop up the network
Using dim_location (location_id, parent_location_id, echelon, location_type) and fct_shipment_leg (leg_id, shipment_id, origin_location_id, destination_location_id, direction, tendered_at_utc, shipped_units), quantify how order variability grows upstream. fct_shipment_leg carries no sku_id and no link back to order lines, so shipped_units is the node's whole mixed-SKU flow and the measurement is node-level by construction. For each node compute weekly units flowing out (what it served downstream) and weekly units flowing in (what it ordered up), then the ratio of their variances on detrended, deseasonalised residuals. Walk parent_location_id from a store to echelon 0 and report the ratio at each hop, naming the hop where amplification is introduced. dim_location has no path column; build the chain yourself.
Approach
- Date orders by tendered_at_utc, not delivered_at_utc. Tender is the closest observable proxy for the moment the ordering decision was made; dating by delivery shifts the series by transit time and smears the variance you are trying to measure.
- Build both series on the same calendar weeks and in the same units. Variance scales with the aggregation window, so a ratio formed from daily inbound against weekly outbound measures the calendar, not the ordering rule. The units are mixed-SKU counts because the leg table carries no sku_id, so a node whose product mix drifts toward smaller or larger pack sizes moves both series for reasons unrelated to its ordering rule; print that caveat next to the number rather than implying a per-SKU result.
- Detrend and deseasonalise before taking variances, or compute the ratio on residuals from a simple weekly seasonal baseline. A growing node otherwise scores as amplifying, when all you have measured is its trend.
- Walk the parent chain iteratively: start from the store rows, join dim_location to itself on parent_location_id, and repeat until parent is null, capping the loop at the known echelon depth and asserting the path length matches the echelon difference so a data cycle raises rather than hangs.
- Read the output as a sequence, not a set of numbers. A pass-through node such as a cross-dock should sit near 1.0, and the hop where the ratio jumps is where a batching rule, a minimum order quantity or truckload rounding lives. That is the node to fix, not the node that is complaining.
Worked solution 45 min
- Assign an ISO week from tendered_at_utc and build two weekly series per node from fct_shipment_leg alone: outbound units summed where origin_location_id is the node, inbound units summed where destination_location_id is the node. No SKU filter is applied because the table has no sku_id, so both series are total unit flow.
- Regress each series on a linear trend plus week-of-year dummies, or subtract a centred moving average, and keep the residuals.
- Compute the amplification ratio per node as var(inbound residuals) / var(outbound residuals) over the shared weeks, requiring a minimum of 26 weeks before reporting a ratio.
- Build the parent chain from a chosen store by iteratively joining dim_location on parent_location_id until parent_location_id is null, asserting the loop terminates within the echelon depth.
- Emit the chain in order with echelon, location_type, both variances, the ratio, the week count and the mixed-SKU caveat, and mark the hop with the largest increase.
Follow-up
- The ratio at one hop is 4.2 and the node insists it orders exactly to forecast. What ordering rules would produce that number anyway?
- What would it take to run this for a single SKU family rather than total units, given fct_shipment_leg carries neither a sku_id nor any link to order lines, and which hops would still be unmeasurable after you added that link?
- What would you expect this measurement to look like during a promotion, and does that change your conclusion?
Write a Python script to merge two large datasets: one containing vehi…
Write a Python script to merge two large datasets: one containing vehicle trip logs and the other containing passenger ride requests, optimizing for memory efficiency.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Explain how you would optimize a slow-running SQL query that aggregate…
Explain how you would optimize a slow-running SQL query that aggregates millions of daily telemetry data points.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
How would you clean and preprocess a highly noisy dataset containing G…
How would you clean and preprocess a highly noisy dataset containing GPS coordinates and timestamps from an autonomous vehicle?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Perfect order rate against the original promise in local time
Compute weekly perfect order rate at order grain. An order qualifies only if every fct_order_line has line_status = 'shipped_full', every line's actual_delivery_at_utc falls on or before original_promised_delivery_date at the ship-from node's local end of day (dim_location.timezone holds IANA names), and no fct_shipment_leg reached through shipment_id carries exception_code in ('damage','address_issue') or leg_status in ('refused','lost'). The denominator is orders whose original_promised_delivery_date lands in the week. Undelivered lines fail. Publish with a fourteen-day lag and declare invoice accuracy as out of scope.
Approach
- Convert the event into local time rather than comparing against a UTC-anchored deadline: (actual_delivery_at_utc AT TIME ZONE 'UTC' AT TIME ZONE loc.timezone)::date <= original_promised_delivery_date. Casting the promise date to a timestamp puts the cut-off at UTC midnight starting the promise day, which is up to a full day early and wrong in both directions depending on the node's offset.
- Evaluate the three components as booleans at line grain, then roll to the order with BOOL_AND, coercing a NULL actual_delivery_at_utc to FALSE explicitly so an undelivered line fails rather than propagating NULL and dropping the order out of the comparison.
- Test the shipment exceptions with EXISTS against fct_shipment_leg, never with a join. One shipment has many legs and one order has many lines, so a join multiplies the order into the denominator and changes the metric.
- Aggregate to order grain first, then count qualifying orders over orders promised in the week. The rate is orders over orders; a line-level average is a different and easier metric.
- Apply the publication lag by restricting to weeks whose promise dates are at least fourteen days old, and emit the three component pass rates beside the headline so any movement is attributable to completeness, timeliness or exceptions.
Worked solution 40 min
- CTE line_flags: join fct_order_line to dim_location on ship_from_location_id, producing is_full, is_on_time using the AT TIME ZONE conversion with COALESCE to FALSE, and has_exception from an EXISTS over fct_shipment_leg.
- CTE order_flags: GROUP BY order_id with BOOL_AND(is_full), BOOL_AND(is_on_time), BOOL_AND(NOT has_exception) and MIN(original_promised_delivery_date) for bucketing.
- Aggregate per DATE_TRUNC('week', promise date): COUNT() as denominator, COUNT() FILTER (WHERE all three) as perfect_orders, and the three component counts.
- Restrict output to weeks ending at least fourteen days ago.
- Re-run the on-time expression once with a naive UTC comparison and record the delta as evidence the conversion is live.
Follow-up
- All three component rates are flat but perfect order fell two points. How is that arithmetically possible, and what does it tell you about where the failures now land?
- A node sits at UTC+13 and many orders are promised on the last day of the month. What does that do to your weekly and monthly buckets?
- How do you handle an order whose lines ship from two nodes in different timezones, and which node's clock owns the promise?
Our autonomous vehicles are experiencing intermittent delays at a spec…
Our autonomous vehicles are experiencing intermittent delays at a specific intersection. How would you use telemetry and operational data to diagnose the root cause?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
How would you prioritize which data features to store and analyze when…
How would you prioritize which data features to store and analyze when faced with bandwidth constraints on vehicle-to-cloud data transmission?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
A planner scorecard that aggregation cannot flatter
Demand planners will be scored on forecast accuracy. fct_inventory_daily carries demand_qty, forecast_qty_lag7, forecast_qty_lag28 and stockout_flag at sku by location by date, with NULL forecasts before a series existed. Specify the scorecard: error metric, grain, lag, exclusion policy and guardrail. Then show the three ways a planner raises the headline without forecasting better, and state which single number you would prefer to accuracy if you could publish only one.
Approach
- Fix the grain and lag to the decision. The ordering policy consumes a forecast at one horizon for one sku-location-week cell, so the score is computed there; a number quoted at any other grain is not the number the policy used and cannot be compared to it.
- Choose WMAPE, sum of absolute errors over sum of actuals, or a pinball loss at the service quantile the policy targets. Rule out MAPE explicitly: it is undefined at zero actuals, which is most days for a C item at a forward-stocking location, and it is asymmetric where it is defined, so optimising it drives systematic under-forecasting.
- Show why aggregation flatters, and show it as arithmetic rather than opinion. The absolute value of a sum is at most the sum of absolute values, so errors cancelling across SKUs, locations or weeks make aggregated WMAPE weakly lower than the disaggregated figure on the same denominator, with no forecasting improvement involved.
- Name the other two routes: scoring at a shorter lag than ordering actually uses, and exclusion drift, where NULL-forecast rows, zero-demand cells or stockout days are dropped without publishing the excluded count, quietly removing the hardest cells from the sample.
- Set the guardrail on signed bias, weighted mean percentage error over the same window and cells, reported per ABC/XYZ cell. A series can be inaccurate but unbiased, or accurate on average and persistently short, and only the second wrecks inventory in one direction.
- Answer the last question directly: prefer the outcome that accuracy is a means to, achieved fill rate at a given inventory at cost per cell, because that is what the business buys, and accuracy that does not move it is not worth paying for.
Worked solution 40 min
- Build the scored panel at sku by location by week, summing demand_qty and taking the forecast at the lag the ordering policy uses, excluding NULL-forecast cells and counting every exclusion.
- Compute WMAPE and signed WMPE per planner per ABC/XYZ cell by re-summing numerator and denominator, never by averaging cell ratios.
- Recompute the same WMAPE at region and at month to demonstrate the mechanical improvement, and record the difference as the aggregation premium.
- Recompute with and without stockout days and with and without zero-demand cells, recording each shift as the exclusion premium.
- Publish WMAPE, WMPE, the excluded-row count and the paired inventory-and-service outcome, with grain and lag printed on the page rather than held in a footnote.
Follow-up
- Two planners have identical WMAPE and opposite bias. Which one is doing more damage, and to what?
- The scorecard is per planner but forecasts are hierarchical and reconciled. Whose number is the reconciled one, and who gets charged for the reconciliation loss?
Perfect order rate falls only in the newest two weeks
A weekly perfect order chart holds at 94 percent for eleven weeks, then reads 91 percent and 84 percent in the two most recent weeks. The metric requires every line shipped_full, actual_delivery_at_utc on or before original_promised_delivery_date, and no joined shipment leg carrying exception_code IN ('damage','address_issue') or leg_status IN ('refused','lost'). It is published on a 14-day lag. Using fct_order_line and fct_shipment_leg, decide whether service regressed and specify what the chart should show for incomplete weeks.
Approach
- Start from the structure of the metric: the denominator is fixed at promise date while the numerator requires an event that may not have happened yet, so any week younger than the full delivery-plus-claims window is mechanically depressed. Establish that before hunting for a cause.
- Quantify the incompleteness instead of assuming it. Per promise week, report the share of lines with actual_delivery_at_utc IS NULL and the share of shipments whose legs are still leg_status IN ('tendered','in_transit').
- Build the maturation curve from history. For each of the eleven settled weeks, recompute the metric as it would have looked at 3, 7 and 14 days after the promise date, so a recent week can be compared against a healthy week at the same age rather than against a settled one.
- Only after age-matching, test whether any residual is real by cutting it by ship_from_location_id, mode and exception_code, and asking whether the loss concentrates where a genuine failure would concentrate.
- Fix the chart rather than only the analysis: suppress or grey weeks younger than the publication lag, or publish an age-matched estimate with an interval, and write the rule down so the next reader is not caught by the same shape.
Follow-up
- Damage claims can arrive up to thirty days later. Does that argue for a longer lag or for a different numerator?
- How would you detect a genuine regression inside the lag window without waiting two weeks for it to settle?
- P90 order cycle time shows the same maturation shape but biased the other way. Why, and which metric is more dangerous to read early?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.
How do you handle missing or corrupted sensor data in a time-series da…
How do you handle missing or corrupted sensor data in a time-series dataset without biasing your downstream models?
Approach
- Pick a story where you drove the decision, not one where you observed it.
- State the situation in two sentences and spend the rest on your reasoning.
- Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
- How did you know the outcome was caused by your change?
- What did you decide not to do, and why?
Walk through an inventory analysis that turned out wrong
Six months ago you published inventory turns by node using the month-end on_hand_qty snapshot from fct_inventory_daily as the denominator, and two DCs cut cover on the strength of it. A finance review later showed month-end is systematically the lowest on-hand point of the month, because shipments cluster before close, so your turns were overstated and days of supply understated. Tell the story: what you built, how the error surfaced, what it cost, what you corrected, and the control that now prevents this class of mistake rather than this specific instance.
Approach
- Give the facts in order and own the decision, not just the code: you chose the month-end snapshot because it was one row per sku-location and fast, and you did not check whether the sampling point was representative.
- Quantify the error rather than describing it: recompute the same window against the average of every daily snapshot and state the gap in turns and in days of supply, which in a network with close-period push typically runs ten to twenty percent.
- Separate the consequence from the mistake honestly: say what the two DCs did, whether service actually degraded, and if it did not, say so instead of inflating the damage to sound accountable.
- Describe the correction and the notification: who was told, how quickly, and whether the restated number changed the recommendation.
- Close on the generalised control: a denominator convention written into the metric definition, plus a row-count assertion at each join grain, because the same shape of error appears when a daily snapshot is joined to shipment events on date equality and one shipment with several legs fans the snapshot out.
Follow-up
- How did you decide whom to tell first, and how did you phrase it?
- What made you trust the month-end snapshot in the first place, and what would have caught it in review?
Disagree with a planning manager about a safety stock cut
A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.
Approach
- Lead with the decision at risk, not the metric error: a 30 percent cut on intermittent items is the cheapest way to convert a measurement artefact into stockouts eight weeks later, by which time nobody will connect the two.
- Name the two defects precisely. Dropping zero-actual rows is not a rounding choice, it removes most of the history for a C-class item at a forward-stocking location and it removes the rows where over-forecasting is penalised, so the surviving MAPE is biased toward whichever series stocked out. Using shipped_qty as the actual scores the forecast against a supply ceiling, so a series looks more accurate the more often it ran out.
- Offer the replacement metric in the same breath: WMAPE, sum of absolute errors over sum of actuals against demand_qty, computed at the sku-location-week grain the ordering decision uses and at the lag it uses, reported next to signed bias so a persistently short series cannot hide inside an accuracy number.
- Propose the smaller action that is defensible now: recompute on the corrected basis, rank cells by bias rather than accuracy, and cut cover only where the corrected series is unbiased or over-forecasting, in a staged rollout with a service guardrail.
- Give him the win he actually wants: if the corrected numbers still support cuts on a subset, say so in advance, so the disagreement is about evidence rather than about territory.
Follow-up
- He says demand_qty is itself incomplete because customers stop ordering what shows out of stock. Is he right, and what do you do about it?
- Which items would you leave alone regardless of what the corrected metric says?
- 01
How do you handle missing or corrupted sensor data in a time-series dataset without biasing your downstream models?
- 02
Six months ago you published inventory turns by node using the month-end on_hand_qty snapshot from fct_inventory_daily as the denominator, and two DCs cut cover on the strength of it. A finance review later showed month-end is systematically the lowest on-hand point of the month, because shipments cluster before close, so your turns were overstated and days of supply understated. Tell the story: what you built, how the error surfaced, what it cost, what you corrected, and the control that now prevents this class of mistake rather than this specific instance.
- 03
A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.
Is this an official May Mobility interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at May Mobility. Rounds and questions reflect what candidates have reported, not a process May Mobility has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long do I have to complete the technical take-home assessment?
You are typically given two to three days to complete the assessment once it is sent to you. The assessment is designed to take approximately 3 to 6 hours of focused work, depending on your familiarity with the dataset style.
PracHub interview research ↗Can I skip the take-home assessment if I have a strong GitHub portfolio?
No. The technical take-home assessment is a mandatory step in the May Mobility interview process for all Data Scientist candidates. It ensures a standardized, objective evaluation of core coding and analytical skills across all applicants.
PracHub interview research ↗Is the work environment fully remote, hybrid, or onsite?
While May Mobility has its headquarters in Ann Arbor, MI, work arrangements depend on the specific team and role requirements. Many data science positions offer hybrid or remote flexibility, but you should clarify expectations with your recruiter during the initial call.
PracHub interview research ↗What is the primary tech stack used by the data science team?
The team primarily utilizes Python for data analysis, modeling, and scripting, alongside SQL for data extraction. Cloud infrastructure is heavily integrated, utilizing modern data warehousing and pipeline tools to manage autonomous vehicle telemetry.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22