Design a text-to-video GPU job scheduler with one GPU per task and fair queueing

Quick Overview

Design a GPU job scheduler for a text-to-video generation service where each task runs on a single GPU and users share capacity through fair queueing. It tests job store design, fairness by GPU time rather than task count, leases and retries, and failure handling across workers, GPUs and the scheduler.

Design a text-to-video GPU job scheduler with one GPU per task and fair queueing

Company: OpenAI

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design the backend of a text-to-video generation service. Users submit generation requests, and each request becomes a task that runs on a GPU fleet. The prompt states two constraints: - Each task runs on a single GPU; it is never split across GPUs. - Scheduling must use fair queueing, so that users share the fleet and no single user can monopolize it. Treat this as a job-scheduling problem. Expect the core of the discussion to be failure handling: what happens when workers, GPUs, tasks or the scheduler itself fail. ### Clarifying Questions - What is the request rate, how long does one task take, and how widely does that duration vary? - What is the unit of fairness: a user, an organization, or a paid tier? Should tiers get different shares? - Can a GPU run more than one task at a time? - Can a running task be preempted, or checkpointed partway through? - How are results delivered: polling, push notification, or both? - Can users cancel a task that is queued or already running? ### Part 1 — Request flow with one GPU per task Design the path from a submitted request to a finished video: the APIs, the job data model, how tasks are matched to GPUs, and where results go. ```hint What the constraint buys and costs One GPU per task removes cross-GPU coordination. Think about what it means for short tasks waiting behind long ones, and what the scheduler must know about each GPU. ``` #### What This Part Should Cover - APIs and a durable job store that is the source of truth - How an idle GPU gets its next task, and how tasks are matched to GPU capabilities - Where outputs are stored and how users get them ### Part 2 — Fair queueing Design the queueing policy that decides which user's task runs next on a free GPU. ```hint Fair in what unit Counting tasks is not the same as counting GPU time when task lengths vary. ``` #### What This Part Should Cover - The fairness metric and the selection rule, including weights for tiers - Treatment of a user who returns after being idle, and per-user concurrency caps - Cost of a scheduling decision, and admission control when demand exceeds capacity ### Part 3 — Failure handling Walk through what can fail between submission and delivery, how each failure is detected, and how the system recovers without losing or duplicating work. ```hint Every step can fail List the failures between submission and delivery. For each one, decide who notices it and what makes the retry safe. ``` #### What This Part Should Cover - Worker and GPU failures, detected through leases and health checks - Tasks that fail repeatedly, tasks that hang, and protection against duplicate results - Failure of the scheduler itself, and recovery of its state ### What a Strong Answer Covers - A durable job store as the source of truth, with the scheduler's queues rebuildable from it - Fairness measured in weighted GPU time, with idle-credit handling and per-user caps - Leases, heartbeats, fencing and bounded retries that make re-execution safe - Handling of long tasks under the one-GPU-per-task constraint - Estimates of fleet throughput and queue growth that inform admission control ### Follow-up Questions - A paid tier must get results faster. How do weights, reserved capacity and preemption compare? - A new model version makes some tasks fail on certain GPU types. How does the system contain the damage? - How would you show each user an honest estimate of their wait time?

Overview: Design a GPU job scheduler for a text-to-video generation service where each task runs on a single GPU and users share capacity through fair queueing. It tests job store design, fairness by GPU time rather than task count, leases and retries, and failure handling across workers, GPUs and the scheduler.

|Home/System Design/OpenAI
OpenAI logo
OpenAI
Sep 28, 2026
mediumSoftware EngineerOnsiteSystem Design
2
0

Design the backend of a text-to-video generation service. Users submit generation requests, and each request becomes a task that runs on a GPU fleet. The prompt states two constraints:

  • Each task runs on a single GPU; it is never split across GPUs.
  • Scheduling must use fair queueing, so that users share the fleet and no single user can monopolize it.

Treat this as a job-scheduling problem. Expect the core of the discussion to be failure handling: what happens when workers, GPUs, tasks or the scheduler itself fail.

Clarifying Questions Guidance

  • What is the request rate, how long does one task take, and how widely does that duration vary?
  • What is the unit of fairness: a user, an organization, or a paid tier? Should tiers get different shares?
  • Can a GPU run more than one task at a time?
  • Can a running task be preempted, or checkpointed partway through?
  • How are results delivered: polling, push notification, or both?
  • Can users cancel a task that is queued or already running?

Part 1 — Request flow with one GPU per task

Design the path from a submitted request to a finished video: the APIs, the job data model, how tasks are matched to GPUs, and where results go.

What This Part Should Cover Guidance

  • APIs and a durable job store that is the source of truth
  • How an idle GPU gets its next task, and how tasks are matched to GPU capabilities
  • Where outputs are stored and how users get them

Part 2 — Fair queueing

Design the queueing policy that decides which user's task runs next on a free GPU.

What This Part Should Cover Guidance

  • The fairness metric and the selection rule, including weights for tiers
  • Treatment of a user who returns after being idle, and per-user concurrency caps
  • Cost of a scheduling decision, and admission control when demand exceeds capacity

Part 3 — Failure handling

Walk through what can fail between submission and delivery, how each failure is detected, and how the system recovers without losing or duplicating work.

What This Part Should Cover Guidance

  • Worker and GPU failures, detected through leases and health checks
  • Tasks that fail repeatedly, tasks that hang, and protection against duplicate results
  • Failure of the scheduler itself, and recovery of its state

What a Strong Answer Covers Guidance

  • A durable job store as the source of truth, with the scheduler's queues rebuildable from it
  • Fairness measured in weighted GPU time, with idle-credit handling and per-user caps
  • Leases, heartbeats, fencing and bounded retries that make re-execution safe
  • Handling of long tasks under the one-GPU-per-task constraint
  • Estimates of fleet throughput and queue growth that inform admission control

Follow-up Questions Guidance

  • A paid tier must get results faster. How do weights, reserved capacity and preemption compare?
  • A new model version makes some tasks fail on certain GPU types. How does the system contain the damage?
  • How would you show each user an honest estimate of their wait time?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...