Applying LLMs to Autonomous Driving Despite Hallucination and Real-Time Limits

Read the full interview experience this question came from →

Quick Overview

An open-ended ML system design discussion on how large language models and vision-language models could be used in autonomous driving, including whether one could actually drive the vehicle. It tests mapping off-board and on-vehicle uses, containing hallucination and latency in a safety-critical stack, and planning evaluation.

Applying LLMs to Autonomous Driving Despite Hallucination and Real-Time Limits

Company: Aurora

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: hard

Interview Round: Technical Screen

You are interviewing for a machine learning engineer role at a company that builds self-driving vehicles. The interviewer keeps returning to one theme: **how would you use large language models (LLMs) in autonomous driving, and could an LLM actually drive a driverless vehicle?** The interviewer knows the standard objection: LLMs hallucinate, and driving is a real-time, safety-critical control problem. They are not looking for a dismissal. They want to see whether you can find where language-model capabilities genuinely help an autonomy program and how you would contain their failure modes. ### Clarifying Questions - Is the question about on-vehicle use (running while the vehicle drives), off-board use (the data, labeling, simulation, and validation pipeline), or both? - What onboard compute and planning-cycle latency budget does the vehicle stack have? If the interviewer leaves this open, state your assumption. - Is the existing driving stack modular (perception, prediction, planning, control) or end-to-end, and is there an independent safety or validation layer that checks planned trajectories? - Does "LLM" include vision-language models that consume camera images, or only text models? - Are rider-facing or remote-operator interactions (explaining behavior, remote assistance) in scope? ### Part 1 — Where LLMs Add Value in an Autonomy Program Map the concrete places where an LLM or vision-language model could improve an autonomous-driving program, both off the vehicle and on it. For each use, state the model's input and output and what it replaces or speeds up. ```hint Look beyond the steering wheel Much of the engineering effort in autonomy goes into finding, labeling, and simulating rare situations, not into the control loop itself. ``` #### What This Part Should Cover - A clear split between off-board and on-vehicle uses, with a risk level for each - Concrete input and output definitions instead of "use an LLM for X" - How each use targets the long tail of rare driving situations - Human review or automated validation wherever model output feeds training or evaluation data ### Part 2 — Could an LLM Drive the Vehicle? Now take the harder version: design how language-model reasoning could influence real driving decisions on the vehicle. Address hallucination, latency, and determinism directly, and spell out what authority the model is and is not given. ```hint Sort decisions by deadline Ask which decisions must be made every planning cycle and which could tolerate a slower opinion that sometimes arrives late or not at all. ``` ```hint Make claims checkable Free-form text is hard to verify. Think about what output format would let the rest of the stack check what the model claims. ``` #### Clarifying Questions for this Part - If the model's output arrives late or not at all, what must the vehicle do? - Is the goal to handle rare, semantically unusual scenes better, or to replace the planner outright? #### What This Part Should Cover - An architecture that keeps a verifiable, deadline-meeting path in control - How hallucinated or stale outputs are detected, bounded, or ignored - A latency and compute strategy for running a large model on the vehicle - A position on end-to-end vision-language-action driving and the problem of validating it ### What a Strong Answer Covers - A constructive stance that takes hallucination seriously without refusing the premise - Safety-case reasoning: what the model's worst output could cause and what bounds it - An evaluation plan (offline benchmarks, log replay, closed-loop simulation, shadow mode) with metrics tied to the targeted scenario classes - Explicit trade-offs between general world knowledge and verifiability - Awareness of new failure modes and attack surfaces that language models introduce ### Follow-up Questions - How would you measure the hallucination rate of a vision-language model on driving scenes, and what rate would be acceptable for each use you proposed? - A road sign or a sticker on a vehicle contains text that reads like an instruction. How does your design stay robust to that kind of injected input? - How would you distill a large model's semantic knowledge into something small enough for the onboard compute budget? - What evidence would convince a safety team to let the model's output influence on-road behavior at all?

Overview: An open-ended ML system design discussion on how large language models and vision-language models could be used in autonomous driving, including whether one could actually drive the vehicle. It tests mapping off-board and on-vehicle uses, containing hallucination and latency in a safety-critical stack, and planning evaluation.

Read the full Aurora Machine Learning Engineer interview experience this question came from

|Home/ML System Design/Aurora
Aurora logo
Aurora
Sep 23, 2026
hardMachine Learning EngineerTechnical ScreenML System Design
0
0

You are interviewing for a machine learning engineer role at a company that builds self-driving vehicles. The interviewer keeps returning to one theme: how would you use large language models (LLMs) in autonomous driving, and could an LLM actually drive a driverless vehicle?

The interviewer knows the standard objection: LLMs hallucinate, and driving is a real-time, safety-critical control problem. They are not looking for a dismissal. They want to see whether you can find where language-model capabilities genuinely help an autonomy program and how you would contain their failure modes.

Clarifying Questions Guidance

  • Is the question about on-vehicle use (running while the vehicle drives), off-board use (the data, labeling, simulation, and validation pipeline), or both?
  • What onboard compute and planning-cycle latency budget does the vehicle stack have? If the interviewer leaves this open, state your assumption.
  • Is the existing driving stack modular (perception, prediction, planning, control) or end-to-end, and is there an independent safety or validation layer that checks planned trajectories?
  • Does "LLM" include vision-language models that consume camera images, or only text models?
  • Are rider-facing or remote-operator interactions (explaining behavior, remote assistance) in scope?

Part 1 — Where LLMs Add Value in an Autonomy Program

Map the concrete places where an LLM or vision-language model could improve an autonomous-driving program, both off the vehicle and on it. For each use, state the model's input and output and what it replaces or speeds up.

What This Part Should Cover Guidance

  • A clear split between off-board and on-vehicle uses, with a risk level for each
  • Concrete input and output definitions instead of "use an LLM for X"
  • How each use targets the long tail of rare driving situations
  • Human review or automated validation wherever model output feeds training or evaluation data

Part 2 — Could an LLM Drive the Vehicle?

Now take the harder version: design how language-model reasoning could influence real driving decisions on the vehicle. Address hallucination, latency, and determinism directly, and spell out what authority the model is and is not given.

Clarifying Questions for this Part Guidance

  • If the model's output arrives late or not at all, what must the vehicle do?
  • Is the goal to handle rare, semantically unusual scenes better, or to replace the planner outright?

What This Part Should Cover Guidance

  • An architecture that keeps a verifiable, deadline-meeting path in control
  • How hallucinated or stale outputs are detected, bounded, or ignored
  • A latency and compute strategy for running a large model on the vehicle
  • A position on end-to-end vision-language-action driving and the problem of validating it

What a Strong Answer Covers Guidance

  • A constructive stance that takes hallucination seriously without refusing the premise
  • Safety-case reasoning: what the model's worst output could cause and what bounds it
  • An evaluation plan (offline benchmarks, log replay, closed-loop simulation, shadow mode) with metrics tied to the targeted scenario classes
  • Explicit trade-offs between general world knowledge and verifiability
  • Awareness of new failure modes and attack surfaces that language models introduce

Follow-up Questions Guidance

  • How would you measure the hallucination rate of a vision-language model on driving scenes, and what rate would be acceptable for each use you proposed?
  • A road sign or a sticker on a vehicle contains text that reads like an instruction. How does your design stay robust to that kind of injected input?
  • How would you distill a large model's semantic knowledge into something small enough for the onboard compute budget?
  • What evidence would convince a safety team to let the model's output influence on-road behavior at all?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...