Where to apply ML in a task flow from customer request to driver-uploaded evidence

Read the full interview experience this question came from →

Quick Overview

An ML system design discussion about a task platform: walk the end-to-end flow from a customer submitting a task to a driver completing it and uploading evidence, and identify where machine learning and LLMs add value. Tests prioritization, model framing, labels, evaluation and risk handling.

Where to apply ML in a task flow from customer request to driver-uploaded evidence

Company: Uber

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Technical Screen

This was a short, free-form design discussion at the end of a phone screen. The interviewer described a platform on which customers submit tasks that drivers on the platform carry out: a customer places a task, a driver completes it, and the driver uploads evidence that the task was done. The question was: **Across this end-to-end flow, from the moment a customer submits a task until a driver has completed it and uploaded evidence, where could machine learning, including large language models, be leveraged?** Map the flow, identify the opportunities, and go deeper on the two or three you consider most valuable: what each model would do, what data and labels it would learn from, how you would evaluate it, and what could go wrong. ```hint Find the human judgment calls Walk a task from the customer's request to the uploaded evidence and mark every step where a person currently reads, writes, decides or checks something by hand. ``` ```hint Labels decide feasibility For each candidate opportunity, ask what a wrong prediction would cost at that step and whether past outcomes already provide labels to train and evaluate a model. ``` ### Constraints and Clarifications - The task types, evidence formats, volumes and current tooling were not specified. Ask about them or state your assumptions. - The discussion is short: cover the whole flow first, then go deep on a few ranked opportunities rather than every idea. ### Clarifying Questions - What kinds of tasks do customers submit, and what counts as evidence: photos, videos, documents, text, location data? - How do customers describe a task today: free text, a form, or a conversation with a sales or operations team? - How is evidence reviewed today, and how many submissions arrive per day? - Which error is costlier: accepting bad evidence, or rejecting a driver's valid work? - Does a rejection of evidence affect the driver's pay for the task? ### What a Strong Answer Covers - A map of the flow from task submission to evidence upload, with each decision point and hand-off named - Opportunities ranked by impact, feasibility and label availability rather than listed - For the top opportunities: the prediction or generation target, the inputs, the labels, and the choice between an LLM and a task-specific model - Offline and online evaluation, with a human-review path where errors are costly - Risks specific to this flow: fraudulent or recycled evidence, fairness and appeals for drivers, privacy of uploaded content, and the cost and latency of model calls ### Follow-up Questions - At launch you have few labeled rejections of evidence. How do you train and evaluate an evidence-quality model? - A driver disputes an automatic rejection. How does your design handle the appeal, and how do appeals feed back into the model? - How would you detect drivers who upload recycled, staged or edited evidence? - Instructions generated by an LLM turn out to be wrong for some new task types. How do you catch that before drivers act on them?

Overview: An ML system design discussion about a task platform: walk the end-to-end flow from a customer submitting a task to a driver completing it and uploading evidence, and identify where machine learning and LLMs add value. Tests prioritization, model framing, labels, evaluation and risk handling.

Read the full Uber Machine Learning Engineer interview experience this question came from

|Home/ML System Design/Uber
Uber logo
Uber
Aug 5, 2026
mediumMachine Learning EngineerTechnical ScreenML System Design
0
0

This was a short, free-form design discussion at the end of a phone screen. The interviewer described a platform on which customers submit tasks that drivers on the platform carry out: a customer places a task, a driver completes it, and the driver uploads evidence that the task was done. The question was:

Across this end-to-end flow, from the moment a customer submits a task until a driver has completed it and uploaded evidence, where could machine learning, including large language models, be leveraged?

Map the flow, identify the opportunities, and go deeper on the two or three you consider most valuable: what each model would do, what data and labels it would learn from, how you would evaluate it, and what could go wrong.

Constraints and Clarifications

  • The task types, evidence formats, volumes and current tooling were not specified. Ask about them or state your assumptions.
  • The discussion is short: cover the whole flow first, then go deep on a few ranked opportunities rather than every idea.

Clarifying Questions Guidance

  • What kinds of tasks do customers submit, and what counts as evidence: photos, videos, documents, text, location data?
  • How do customers describe a task today: free text, a form, or a conversation with a sales or operations team?
  • How is evidence reviewed today, and how many submissions arrive per day?
  • Which error is costlier: accepting bad evidence, or rejecting a driver's valid work?
  • Does a rejection of evidence affect the driver's pay for the task?

What a Strong Answer Covers Guidance

  • A map of the flow from task submission to evidence upload, with each decision point and hand-off named
  • Opportunities ranked by impact, feasibility and label availability rather than listed
  • For the top opportunities: the prediction or generation target, the inputs, the labels, and the choice between an LLM and a task-specific model
  • Offline and online evaluation, with a human-review path where errors are costly
  • Risks specific to this flow: fraudulent or recycled evidence, fairness and appeals for drivers, privacy of uploaded content, and the cost and latency of model calls

Follow-up Questions Guidance

  • At launch you have few labeled rejections of evidence. How do you train and evaluate an evidence-quality model?
  • A driver disputes an automatic rejection. How does your design handle the appeal, and how do appeals feed back into the model?
  • How would you detect drivers who upload recycled, staged or edited evidence?
  • Instructions generated by an LLM turn out to be wrong for some new task types. How do you catch that before drivers act on them?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...