Design a Scheduler for Weekly Self-Driving Simulation Jobs with Dependent Phases

Read the full interview experience this question came from →

Quick Overview

A system design question about scheduling a weekly batch of about 10,000 self-driving car simulation cases, where each case runs several phases that depend on one another. It tests dependency tracking, dispatch to GPU and CPU worker pools, retries and leases, prioritization against a deadline, and observability.

Design a Scheduler for Weekly Self-Driving Simulation Jobs with Dependent Phases

Company: Nuro

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design a job scheduler for a self-driving car simulation platform. Every week the platform runs about 10,000 simulation cases. Each case is a job made of several phases, and the phases depend on one another: a phase may start only after the phases it depends on have finished. The system accepts the weekly run, schedules every phase onto compute, tracks progress, handles failures, and reports the results. ```hint Find the real bottleneck Turn 10,000 cases a week into phase executions per second and into compute-hours, using assumed phase counts and durations. Compare the two numbers before you choose any infrastructure. ``` ```hint Where dependency state lives Decide which component knows that a phase has finished, and what must happen in that same step so that the phases depending on it become runnable exactly once. ``` ```hint Workers fail mid-phase A simulation machine can crash or be reclaimed halfway through a phase. Decide how the system notices, what it reruns, and how it ignores a late report from a machine it has already given up on. ``` ### Constraints and Clarifications - About 10,000 simulation cases run each week, and each case has several phases with dependencies between them. - Assume a phase's dependencies are other phases and never form a cycle. - Phase durations, resource needs and the deadline for the weekly run are not given; agree on them with the interviewer (see below). ### Clarifying Questions - Are dependencies only between phases of the same case, or can a phase depend on other cases, for example a report that summarizes the whole run? - Do all 10,000 cases arrive at once as one weekly batch, and by when must the run finish? - How long does each phase take, and which phases need GPUs and which only CPUs? - Do all cases share one phase graph, or can each case define its own? - When a simulation shows the driving software failing a scenario, is that a result to record or an error to retry? - Do other workloads, such as engineers' ad-hoc runs, compete for the same compute, and what priority does the weekly run get? - Must reruns be reproducible, with the same software build, inputs and random seeds? ### What a Strong Answer Covers - Sizing that separates the scheduler's own load from the compute demand, and an explicit completion target for the weekly run - A phase dependency graph with durable per-phase state and an atomic "mark done and unlock dependents" step - Dispatch to heterogeneous worker pools with leases, heartbeats and protection against duplicate or stale completions - A retry policy that separates infrastructure failures from genuine simulation failures, and what happens to downstream phases - Prioritization of the weekly run against other work, attention to the critical path, and autoscaling toward the deadline - Passing artifacts between phases, reproducibility, progress reporting and observability ### Follow-up Questions - A bad simulator build makes every case fail in its first phase. How does the system avoid burning the week's compute budget? - Engineers fix a bug and want to rerun only the cases that failed. What does the system need to support that cheaply? - How would you guarantee the weekly run meets its deadline when ad-hoc runs compete for the same GPUs? - How would you make long simulation phases survive preemption on cheaper, reclaimable machines?

Overview: A system design question about scheduling a weekly batch of about 10,000 self-driving car simulation cases, where each case runs several phases that depend on one another. It tests dependency tracking, dispatch to GPU and CPU worker pools, retries and leases, prioritization against a deadline, and observability.

Read the full Nuro Software Engineer interview experience this question came from

|Home/System Design/Nuro
Nuro logo
Nuro
Aug 21, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Design a job scheduler for a self-driving car simulation platform. Every week the platform runs about 10,000 simulation cases. Each case is a job made of several phases, and the phases depend on one another: a phase may start only after the phases it depends on have finished. The system accepts the weekly run, schedules every phase onto compute, tracks progress, handles failures, and reports the results.

Constraints and Clarifications

  • About 10,000 simulation cases run each week, and each case has several phases with dependencies between them.
  • Assume a phase's dependencies are other phases and never form a cycle.
  • Phase durations, resource needs and the deadline for the weekly run are not given; agree on them with the interviewer (see below).

Clarifying Questions Guidance

  • Are dependencies only between phases of the same case, or can a phase depend on other cases, for example a report that summarizes the whole run?
  • Do all 10,000 cases arrive at once as one weekly batch, and by when must the run finish?
  • How long does each phase take, and which phases need GPUs and which only CPUs?
  • Do all cases share one phase graph, or can each case define its own?
  • When a simulation shows the driving software failing a scenario, is that a result to record or an error to retry?
  • Do other workloads, such as engineers' ad-hoc runs, compete for the same compute, and what priority does the weekly run get?
  • Must reruns be reproducible, with the same software build, inputs and random seeds?

What a Strong Answer Covers Guidance

  • Sizing that separates the scheduler's own load from the compute demand, and an explicit completion target for the weekly run
  • A phase dependency graph with durable per-phase state and an atomic "mark done and unlock dependents" step
  • Dispatch to heterogeneous worker pools with leases, heartbeats and protection against duplicate or stale completions
  • A retry policy that separates infrastructure failures from genuine simulation failures, and what happens to downstream phases
  • Prioritization of the weekly run against other work, attention to the critical path, and autoscaling toward the deadline
  • Passing artifacts between phases, reproducibility, progress reporting and observability

Follow-up Questions Guidance

  • A bad simulator build makes every case fail in its first phase. How does the system avoid burning the week's compute budget?
  • Engineers fix a bug and want to rerun only the cases that failed. What does the system need to support that cheaply?
  • How would you guarantee the weekly run meets its deadline when ad-hoc runs compete for the same GPUs?
  • How would you make long simulation phases survive preemption on cheaper, reclaimable machines?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...