May Mobility · Data Scientist
Updated · 2026-09-22

May Mobility Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

A Data Scientist at May Mobility plays a pivotal role in shaping the future of autonomous vehicle (AV) technology. Operating at the intersection of cutting-edge robotics, machine learning, and transportation logistics, you will transform massive streams of raw vehicle telemetry, sensor logs, and operational data into actionable insights. Your work directly influences how autonomous shuttles navigate complex urban environments, optimize their routes, and ensure passenger safety.

Product-sense cases reward reasoning from a mechanism to a testable prediction. Reciting every metric you can name reads as pattern matching; naming the single quantity that would move if your explanation were true reads as thinking.

May Mobility candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Evaluate forecasts at the lag ordering actually usesJoin daily snapshots to shipment events safelyTrace an on-time miss to one node

34 min read

Practice 15 Data Scientist prompts
15Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Data Scientist at May Mobility plays a pivotal role in shaping the future of autonomous vehicle (AV) technology. Operating at the intersection of cutting-edge robotics, machine learning, and transportation logistics, you will transform massive streams of raw vehicle telemetry, sensor logs, and operational data into actionable insights. Your work directly influences how autonomous shuttles navigate complex urban environments, optimize their routes, and ensure passenger safety.

The impact of this role is profound. Because May Mobility is in an active research and development (R&D) phase, you will not simply maintain existing pipelines; you will design foundational metrics and analytical frameworks from scratch. Whether you are analyzing rider demand patterns, optimizing fleet deployment, or defining safety-critical performance indicators, your models and analyses will drive strategic decisions across engineering, product, and operations teams.

This position is ideal for those who thrive in high-ambiguity environments. You will work with complex, unstructured datasets that represent real-world physical interactions. If you are passionate about applying statistical rigor to physical-world problems and want to see your algorithms directly impact autonomous fleets on public roads, the role offers an unparalleled opportunity for technical ownership and visible real-world impact.

01

Recruiter Screen

reported

A screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.

What to demonstrate

  • Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
  • Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
  • Whether timeline, location and compensation expectations make the rest of the loop worth scheduling

How to prepare

  • Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
  • Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
  • Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
PracHub interview research ↗
02

Technical Take-Home Assessment

reported

A take-home is graded as an argument, not as a notebook. Somebody reads the submission without you in the room, so every choice has to survive on the page: why the question was framed this way, and what was deliberately left out. The gap between a strong and a weak submission is almost never model quality. It is whether the writeup names the specific question it answers and commits to a recommendation, including what evidence would overturn it. A high-accuracy model attached to no conclusion reads as effort that stopped before the decision.

What to demonstrate

  • Whether the question you answered is stated outright, and whether it is the question the prompt posed rather than an easier neighbour of it
  • Whether the recommendation is specific enough to act on, with the uncertainty attached to it instead of parked in a caveats section at the end
  • Whether analytical choices such as the metric definition, the population filter and the time window are justified in the prose, not merely visible in code

How to prepare

  • Take a dataset you have already worked with, write the one-paragraph conclusion first, then check whether the analysis you were planning actually supports it and cut whatever does not
  • Practise stating a metric in one sentence that fixes the population, the time window and the denominator, then confirm your query computes exactly that sentence and nothing adjacent to it
  • Hand a draft to someone outside the problem and ask them to tell you back what you recommended and why; anything they cannot recover is not on the page yet
PracHub interview research ↗
03

Interviews with Hiring Manager

reported

Much of this round runs on your own history, but the manager is not collecting a project list. They are working out what it is like when something goes wrong on your watch: how late the bad news tends to arrive, and whether a number you hand over has been checked by anyone including you. That is why the strongest material is a project where you can describe the part that did not work and what it cost. A result you cannot take full responsibility for, however clean, gives them nothing to trust you with afterwards.

What to demonstrate

  • Whether you volunteer the limits of a result you are proud of, or wait to be pushed onto them
  • How errors surfaced in your past work, and whether you or somebody else found them
  • Whether the scope you claim matches the level of detail you can still produce about it
  • What you did the first time a stakeholder acted on something of yours that turned out to be wrong

How to prepare

  • Rebuild one headline figure from memory down to the join and the filter, so a question about the denominator does not stall the conversation
  • For each project you raise, write the sentence you would say to someone who had already acted on a number that later turned out wrong
  • Mark which parts of a project were yours and which belonged to other people, and state that boundary yourself before anyone asks
PracHub interview research ↗
04

Interviews with Data Science Team

reported

Because the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.

What to demonstrate

  • Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
  • Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
  • Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs

How to prepare

  • For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
  • Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
  • For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Averaging rates across SKU-locations instead of re-summing

Fill rate, turns, OEE and on-time rate are all ratios whose denominators differ by orders of magnitude between cells, so an unweighted mean gives a slow-moving C item at a small node the same vote as a high-volume A item at a national node. The blended figure then moves whenever the portfolio mix moves, and it can improve in every cell while the company-level ratio worsens, or the reverse, which is Simpson's paradox with a warehouse attached. Always sum numerator and denominator to the reporting level and divide once, and when a rate must be compared across nodes, standardise on a fixed SKU mix before reading anything into the difference.

02

Sizing safety stock as z times sigma_D times the square root of lead time

That form assumes lead time is deterministic. When lead time itself varies, the standard deviation of demand over lead time is sqrt(L_bar * sigma_D^2 + D_bar^2 * sigma_L^2), and the second term dominates whenever supply is unreliable, so the familiar formula can understate the requirement severalfold for a long, variable inbound lane. Two further preconditions are routinely forgotten: demand is assumed independent across periods, which promotions and order batching break, and z maps to cycle service level (the probability of no stockout in a replenishment cycle), not to fill rate, which additionally depends on order quantity through the unit normal loss function. Quoting a z-derived number as a fill rate overstates achieved service, and the gap widens as order quantity shrinks.

03

Generalising beyond the population the sample actually supports

State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.

04

SQL that silently fans out on a one-to-many join

State the grain of each table and the grain you want in the result before writing the join. Pre-aggregate the many side to the join key, or use EXISTS or a window function, and verify with a row count against COUNT(DISTINCT id) rather than trusting that the numbers look plausible.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

12 technical prompts3 include a worked solution

How would you define and measure "ride comfort" using vehicle accelero…

medium
statistics and probability

How would you define and measure "ride comfort" using vehicle accelerometer and gyroscope data?

Approach
  1. Say what the estimate is of, and over what population it generalises.
  2. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  3. Write down the assumption the method needs before you use the method.
Follow-up
  • How would you explain this result to someone who does not know statistics?
  • Which assumption here is most likely to be violated in practice?

How do you determine if a sample size is statistically significant whe…

medium
statistics and probability

How do you determine if a sample size is statistically significant when testing autonomous shuttle performance in a new, low-traffic market?

Approach
  1. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  2. Say what the estimate is of, and over what population it generalises.
  3. Sanity-check the answer against a simple bound or a simulated case.
Follow-up
  • What sample size would you need to detect an effect half this size?
  • How would you explain this result to someone who does not know statistics?

If we introduce a new routing algorithm in a specific city, how would …

medium
machine learning and modelling

If we introduce a new routing algorithm in a specific city, how would you design an A/B test to measure its impact on passenger wait times?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Check what information would not exist at prediction time, and exclude it.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • How would you choose the decision threshold, and who owns that choice?
  • Where could label leakage enter this setup?

Measure demand amplification hop by hop up the network

hardWorked solution
bullwhiphierarchy traversalvariance

Using dim_location (location_id, parent_location_id, echelon, location_type) and fct_shipment_leg (leg_id, shipment_id, origin_location_id, destination_location_id, direction, tendered_at_utc, shipped_units), quantify how order variability grows upstream. fct_shipment_leg carries no sku_id and no link back to order lines, so shipped_units is the node's whole mixed-SKU flow and the measurement is node-level by construction. For each node compute weekly units flowing out (what it served downstream) and weekly units flowing in (what it ordered up), then the ratio of their variances on detrended, deseasonalised residuals. Walk parent_location_id from a store to echelon 0 and report the ratio at each hop, naming the hop where amplification is introduced. dim_location has no path column; build the chain yourself.

Approach
  1. Date orders by tendered_at_utc, not delivered_at_utc. Tender is the closest observable proxy for the moment the ordering decision was made; dating by delivery shifts the series by transit time and smears the variance you are trying to measure.
  2. Build both series on the same calendar weeks and in the same units. Variance scales with the aggregation window, so a ratio formed from daily inbound against weekly outbound measures the calendar, not the ordering rule. The units are mixed-SKU counts because the leg table carries no sku_id, so a node whose product mix drifts toward smaller or larger pack sizes moves both series for reasons unrelated to its ordering rule; print that caveat next to the number rather than implying a per-SKU result.
  3. Detrend and deseasonalise before taking variances, or compute the ratio on residuals from a simple weekly seasonal baseline. A growing node otherwise scores as amplifying, when all you have measured is its trend.
  4. Walk the parent chain iteratively: start from the store rows, join dim_location to itself on parent_location_id, and repeat until parent is null, capping the loop at the known echelon depth and asserting the path length matches the echelon difference so a data cycle raises rather than hangs.
  5. Read the output as a sequence, not a set of numbers. A pass-through node such as a cross-dock should sit near 1.0, and the hop where the ratio jumps is where a batching rule, a minimum order quantity or truckload rounding lives. That is the node to fix, not the node that is complaining.
Worked solution 45 min
  1. Assign an ISO week from tendered_at_utc and build two weekly series per node from fct_shipment_leg alone: outbound units summed where origin_location_id is the node, inbound units summed where destination_location_id is the node. No SKU filter is applied because the table has no sku_id, so both series are total unit flow.
  2. Regress each series on a linear trend plus week-of-year dummies, or subtract a centred moving average, and keep the residuals.
  3. Compute the amplification ratio per node as var(inbound residuals) / var(outbound residuals) over the shared weeks, requiring a minimum of 26 weeks before reporting a ratio.
  4. Build the parent chain from a chosen store by iteratively joining dim_location on parent_location_id until parent_location_id is null, asserting the loop terminates within the echelon depth.
  5. Emit the chain in order with echelon, location_type, both variances, the ratio, the week count and the mixed-SKU caveat, and mark the hop with the largest increase.
EXPECTED RESULTAn ordered table from store to echelon 0, one row per hop, each with the two residual variances, the ratio, the number of weeks it rests on, and a stated caveat that units are pooled across SKUs. Cross-dock and pass-through nodes sit near 1.0; at least one hop should sit materially above 1.0 and that hop is the finding.
Follow-up
  • The ratio at one hop is 4.2 and the node insists it orders exactly to forecast. What ordering rules would produce that number anyway?
  • What would it take to run this for a single SKU family rather than total units, given fct_shipment_leg carries neither a sku_id nor any link to order lines, and which hops would still be unmeasurable after you added that link?
  • What would you expect this measurement to look like during a promotion, and does that change your conclusion?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.

How do you handle missing or corrupted sensor data in a time-series da…

medium
behavioural and stakeholder questions

How do you handle missing or corrupted sensor data in a time-series dataset without biasing your downstream models?

Approach
  1. Pick a story where you drove the decision, not one where you observed it.
  2. State the situation in two sentences and spend the rest on your reasoning.
  3. Name the disagreement or constraint, and how you resolved it with evidence.
Follow-up
  • How did you know the outcome was caused by your change?
  • What did you decide not to do, and why?

Walk through an inventory analysis that turned out wrong

medium
error postmortemsnapshot biasinventory turns

Six months ago you published inventory turns by node using the month-end on_hand_qty snapshot from fct_inventory_daily as the denominator, and two DCs cut cover on the strength of it. A finance review later showed month-end is systematically the lowest on-hand point of the month, because shipments cluster before close, so your turns were overstated and days of supply understated. Tell the story: what you built, how the error surfaced, what it cost, what you corrected, and the control that now prevents this class of mistake rather than this specific instance.

Approach
  1. Give the facts in order and own the decision, not just the code: you chose the month-end snapshot because it was one row per sku-location and fast, and you did not check whether the sampling point was representative.
  2. Quantify the error rather than describing it: recompute the same window against the average of every daily snapshot and state the gap in turns and in days of supply, which in a network with close-period push typically runs ten to twenty percent.
  3. Separate the consequence from the mistake honestly: say what the two DCs did, whether service actually degraded, and if it did not, say so instead of inflating the damage to sound accountable.
  4. Describe the correction and the notification: who was told, how quickly, and whether the restated number changed the recommendation.
  5. Close on the generalised control: a denominator convention written into the metric definition, plus a row-count assertion at each join grain, because the same shape of error appears when a daily snapshot is joined to shipment events on date equality and one shipment with several legs fans the snapshot out.
Follow-up
  • How did you decide whom to tell first, and how did you phrase it?
  • What made you trust the month-end snapshot in the first place, and what would have caught it in review?

Disagree with a planning manager about a safety stock cut

medium
disagreeing with datacensoringforecast evaluation

A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.

Approach
  1. Lead with the decision at risk, not the metric error: a 30 percent cut on intermittent items is the cheapest way to convert a measurement artefact into stockouts eight weeks later, by which time nobody will connect the two.
  2. Name the two defects precisely. Dropping zero-actual rows is not a rounding choice, it removes most of the history for a C-class item at a forward-stocking location and it removes the rows where over-forecasting is penalised, so the surviving MAPE is biased toward whichever series stocked out. Using shipped_qty as the actual scores the forecast against a supply ceiling, so a series looks more accurate the more often it ran out.
  3. Offer the replacement metric in the same breath: WMAPE, sum of absolute errors over sum of actuals against demand_qty, computed at the sku-location-week grain the ordering decision uses and at the lag it uses, reported next to signed bias so a persistently short series cannot hide inside an accuracy number.
  4. Propose the smaller action that is defensible now: recompute on the corrected basis, rank cells by bias rather than accuracy, and cut cover only where the corrected series is unbiased or over-forecasting, in a staged rollout with a service guardrail.
  5. Give him the win he actually wants: if the corrected numbers still support cuts on a subset, say so in advance, so the disagreement is about evidence rather than about territory.
Follow-up
  • He says demand_qty is itself incomplete because customers stop ordering what shows out of stock. Is he right, and what do you do about it?
  • Which items would you leave alone regardless of what the corrected metric says?
  • 01

    How do you handle missing or corrupted sensor data in a time-series dataset without biasing your downstream models?

  • 02

    Six months ago you published inventory turns by node using the month-end on_hand_qty snapshot from fct_inventory_daily as the denominator, and two DCs cut cover on the strength of it. A finance review later showed month-end is systematically the lowest on-hand point of the month, because shipments cluster before close, so your turns were overstated and days of supply understated. Tell the story: what you built, how the error surfaced, what it cost, what you corrected, and the control that now prevents this class of mistake rather than this specific instance.

  • 03

    A planning product manager proposes cutting safety stock 30 percent on every C and Z class sku-location cell, citing a MAPE improvement from 41 to 28 percent over two quarters. Checking the query, you find MAPE is computed over fct_inventory_daily rows where the actual is greater than zero, and the actual used is shipped_qty rather than demand_qty; stockout_flag is true on 14 percent of the rows that were kept. You have fifteen minutes in his planning review. Make the disagreement, and propose what you would do instead of the flat cut.

PracHub interview preparation framework ↗
Is this an official May Mobility interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at May Mobility. Rounds and questions reflect what candidates have reported, not a process May Mobility has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How long do I have to complete the technical take-home assessment?

You are typically given two to three days to complete the assessment once it is sent to you. The assessment is designed to take approximately 3 to 6 hours of focused work, depending on your familiarity with the dataset style.

PracHub interview research ↗
Can I skip the take-home assessment if I have a strong GitHub portfolio?

No. The technical take-home assessment is a mandatory step in the May Mobility interview process for all Data Scientist candidates. It ensures a standardized, objective evaluation of core coding and analytical skills across all applicants.

PracHub interview research ↗
Is the work environment fully remote, hybrid, or onsite?

While May Mobility has its headquarters in Ann Arbor, MI, work arrangements depend on the specific team and role requirements. Many data science positions offer hybrid or remote flexibility, but you should clarify expectations with your recruiter during the initial call.

PracHub interview research ↗
What is the primary tech stack used by the data science team?

The team primarily utilizes Python for data analysis, modeling, and scripting, alongside SQL for data extraction. Cloud infrastructure is heavily integrated, utilizing modern data warehousing and pipeline tools to manage autonomous vehicle telemetry.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.