Design a Distributed Job Scheduler for One-Off and Recurring Jobs
Company: Fluidstack
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design a job scheduler: a service that accepts jobs from clients, runs each one on a pool of worker machines at the right time, and tracks every run until it finishes. Assume the usual scope unless the interviewer narrows it:
- A job can be submitted to run immediately, once at a given future time, or repeatedly on a schedule such as a cron expression.
- Each execution of a job, called a run, has a status that clients can query, along with its outcome and logs.
- Failed runs are retried according to a per-job policy. Clients can pause, update or delete a job, and cancel a run.
```hint Finding what is due
Scanning every job definition every second does not scale to millions of jobs. Think about what to store and index so that each scheduling pass touches only the jobs that are actually due.
```
```hint Everything can fail twice
You will run several scheduler instances and many workers, and any of them can die or stall at any moment. Decide what stops one scheduled run from being launched twice, and what happens to a run whose worker disappears halfway through.
```
### Constraints and Clarifications
- Treat a job as an opaque unit of work, for example a container image or a command with arguments. The scheduler decides when and where it runs, not what it does.
- Scale, start-time precision, job durations and delivery guarantees are not given. Ask for them or state your assumptions.
### Clarifying Questions
- Is the scheduler deciding when jobs run (one-off and recurring triggers), where they run (placing jobs that need CPU, memory or GPUs onto machines), or both?
- How many job definitions and runs per day, and how bursty is the load, for example many jobs scheduled for the top of the hour or for midnight?
- How close to its scheduled time must a run start: within a second, or within a minute?
- Must a run execute at most once, at least once, or effectively once? Can job owners make their jobs idempotent?
- How long do jobs run, from seconds to days? Do they need timeouts and cancellation?
- Are there priorities, per-team quotas, or dependencies between jobs?
- If the scheduler was down when a recurring run was due, should that run happen late, be skipped, or be replayed for every missed slot?
### What a Strong Answer Covers
- A data model that separates job definitions from individual runs, with an explicit run state machine
- An efficient, partitioned way to find due jobs, and how scheduler instances share that work
- How each run is launched once, how lost workers and stuck runs are detected, and what guarantee the system really offers
- Retries with backoff, timeouts, cancellation, and a defined policy for missed recurring runs
- Behavior under bursts and contention: queues, priorities, per-tenant limits and back-pressure
- Failure handling for every component, and the metrics that show the scheduler is healthy
### Follow-up Questions
- One million jobs are scheduled for exactly midnight. What happens in your design, and how do you keep start delays acceptable?
- A worker stops sending heartbeats two hours into a three-hour job. How do you decide whether to rerun it, and what if that worker is still running but cut off from the network?
- How would you add dependencies, so that a job starts only after the jobs it depends on have succeeded?
- If jobs declare CPU, memory or GPU requirements, how does dispatching change?
Overview: A system design question asking you to design a distributed job scheduler that runs jobs immediately, at a set time or on a recurring cron schedule across a pool of workers. It tests how you find due jobs at scale, prevent duplicate runs, recover from worker failures, and handle retries, bursts and missed runs.