Applying LLMs to Autonomous Driving Despite Hallucination and Real-Time Limits
Company: Aurora
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Technical Screen
You are interviewing for a machine learning engineer role at a company that builds self-driving vehicles. The interviewer keeps returning to one theme: **how would you use large language models (LLMs) in autonomous driving, and could an LLM actually drive a driverless vehicle?**
The interviewer knows the standard objection: LLMs hallucinate, and driving is a real-time, safety-critical control problem. They are not looking for a dismissal. They want to see whether you can find where language-model capabilities genuinely help an autonomy program and how you would contain their failure modes.
### Clarifying Questions
- Is the question about on-vehicle use (running while the vehicle drives), off-board use (the data, labeling, simulation, and validation pipeline), or both?
- What onboard compute and planning-cycle latency budget does the vehicle stack have? If the interviewer leaves this open, state your assumption.
- Is the existing driving stack modular (perception, prediction, planning, control) or end-to-end, and is there an independent safety or validation layer that checks planned trajectories?
- Does "LLM" include vision-language models that consume camera images, or only text models?
- Are rider-facing or remote-operator interactions (explaining behavior, remote assistance) in scope?
### Part 1 — Where LLMs Add Value in an Autonomy Program
Map the concrete places where an LLM or vision-language model could improve an autonomous-driving program, both off the vehicle and on it. For each use, state the model's input and output and what it replaces or speeds up.
```hint Look beyond the steering wheel
Much of the engineering effort in autonomy goes into finding, labeling, and simulating rare situations, not into the control loop itself.
```
#### What This Part Should Cover
- A clear split between off-board and on-vehicle uses, with a risk level for each
- Concrete input and output definitions instead of "use an LLM for X"
- How each use targets the long tail of rare driving situations
- Human review or automated validation wherever model output feeds training or evaluation data
### Part 2 — Could an LLM Drive the Vehicle?
Now take the harder version: design how language-model reasoning could influence real driving decisions on the vehicle. Address hallucination, latency, and determinism directly, and spell out what authority the model is and is not given.
```hint Sort decisions by deadline
Ask which decisions must be made every planning cycle and which could tolerate a slower opinion that sometimes arrives late or not at all.
```
```hint Make claims checkable
Free-form text is hard to verify. Think about what output format would let the rest of the stack check what the model claims.
```
#### Clarifying Questions for this Part
- If the model's output arrives late or not at all, what must the vehicle do?
- Is the goal to handle rare, semantically unusual scenes better, or to replace the planner outright?
#### What This Part Should Cover
- An architecture that keeps a verifiable, deadline-meeting path in control
- How hallucinated or stale outputs are detected, bounded, or ignored
- A latency and compute strategy for running a large model on the vehicle
- A position on end-to-end vision-language-action driving and the problem of validating it
### What a Strong Answer Covers
- A constructive stance that takes hallucination seriously without refusing the premise
- Safety-case reasoning: what the model's worst output could cause and what bounds it
- An evaluation plan (offline benchmarks, log replay, closed-loop simulation, shadow mode) with metrics tied to the targeted scenario classes
- Explicit trade-offs between general world knowledge and verifiability
- Awareness of new failure modes and attack surfaces that language models introduce
### Follow-up Questions
- How would you measure the hallucination rate of a vision-language model on driving scenes, and what rate would be acceptable for each use you proposed?
- A road sign or a sticker on a vehicle contains text that reads like an instruction. How does your design stay robust to that kind of injected input?
- How would you distill a large model's semantic knowledge into something small enough for the onboard compute budget?
- What evidence would convince a safety team to let the model's output influence on-road behavior at all?
Overview: An open-ended ML system design discussion on how large language models and vision-language models could be used in autonomous driving, including whether one could actually drive the vehicle. It tests mapping off-board and on-vehicle uses, containing hallucination and latency in a safety-critical stack, and planning evaluation.
Read the full Aurora Machine Learning Engineer interview experience this question came from