Design a text-to-video GPU job scheduler with one GPU per task and fair queueing
Company: OpenAI
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design the backend of a text-to-video generation service. Users submit generation requests, and each request becomes a task that runs on a GPU fleet. The prompt states two constraints:
- Each task runs on a single GPU; it is never split across GPUs.
- Scheduling must use fair queueing, so that users share the fleet and no single user can monopolize it.
Treat this as a job-scheduling problem. Expect the core of the discussion to be failure handling: what happens when workers, GPUs, tasks or the scheduler itself fail.
### Clarifying Questions
- What is the request rate, how long does one task take, and how widely does that duration vary?
- What is the unit of fairness: a user, an organization, or a paid tier? Should tiers get different shares?
- Can a GPU run more than one task at a time?
- Can a running task be preempted, or checkpointed partway through?
- How are results delivered: polling, push notification, or both?
- Can users cancel a task that is queued or already running?
### Part 1 — Request flow with one GPU per task
Design the path from a submitted request to a finished video: the APIs, the job data model, how tasks are matched to GPUs, and where results go.
```hint What the constraint buys and costs
One GPU per task removes cross-GPU coordination. Think about what it means for short tasks waiting behind long ones, and what the scheduler must know about each GPU.
```
#### What This Part Should Cover
- APIs and a durable job store that is the source of truth
- How an idle GPU gets its next task, and how tasks are matched to GPU capabilities
- Where outputs are stored and how users get them
### Part 2 — Fair queueing
Design the queueing policy that decides which user's task runs next on a free GPU.
```hint Fair in what unit
Counting tasks is not the same as counting GPU time when task lengths vary.
```
#### What This Part Should Cover
- The fairness metric and the selection rule, including weights for tiers
- Treatment of a user who returns after being idle, and per-user concurrency caps
- Cost of a scheduling decision, and admission control when demand exceeds capacity
### Part 3 — Failure handling
Walk through what can fail between submission and delivery, how each failure is detected, and how the system recovers without losing or duplicating work.
```hint Every step can fail
List the failures between submission and delivery. For each one, decide who notices it and what makes the retry safe.
```
#### What This Part Should Cover
- Worker and GPU failures, detected through leases and health checks
- Tasks that fail repeatedly, tasks that hang, and protection against duplicate results
- Failure of the scheduler itself, and recovery of its state
### What a Strong Answer Covers
- A durable job store as the source of truth, with the scheduler's queues rebuildable from it
- Fairness measured in weighted GPU time, with idle-credit handling and per-user caps
- Leases, heartbeats, fencing and bounded retries that make re-execution safe
- Handling of long tasks under the one-GPU-per-task constraint
- Estimates of fleet throughput and queue growth that inform admission control
### Follow-up Questions
- A paid tier must get results faster. How do weights, reserved capacity and preemption compare?
- A new model version makes some tasks fail on certain GPU types. How does the system contain the damage?
- How would you show each user an honest estimate of their wait time?
Overview: Design a GPU job scheduler for a text-to-video generation service where each task runs on a single GPU and users share capacity through fair queueing. It tests job store design, fairness by GPU time rather than task count, leases and retries, and failure handling across workers, GPUs and the scheduler.