Design a Reliable Distributed Job Scheduler
Company: Robinhood
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
# Design a Reliable Distributed Job Scheduler
Design a service that accepts jobs to run once at a specified time or repeatedly on a schedule. Clients must be able to create, inspect, pause, resume, and cancel jobs. Multiple workers execute jobs, and the system must continue operating when an individual scheduler or worker fails.
Focus on the scheduling, claiming, execution, and recovery path. Define the delivery guarantee you offer to job handlers and explain what clients must do when exactly-once side effects cannot be guaranteed end to end.
### Constraints & Assumptions
- The system serves many independent tenants and should prevent one tenant from monopolizing workers.
- A job may take longer than expected, a worker may disappear mid-run, and duplicate delivery is possible.
- Schedules and job definitions must survive service restarts.
- You may choose reasonable scale targets after stating them; do not assume a single machine is sufficient.
### Clarifying Questions to Ask
- Are schedules one-time, fixed-delay, or calendar-based, and which time-zone semantics apply?
- What lateness is acceptable, and are jobs allowed to overlap with their previous run?
- Are payloads small references or arbitrary blobs?
- What retry, priority, ordering, and tenant-isolation guarantees matter?
### What a Strong Answer Covers
- Durable job and run models with explicit lifecycle states
- Efficient discovery and partitioning of due work
- Lease-based claiming, heartbeats, retry policy, and dead-letter handling
- Idempotency guidance and a candid delivery-guarantee discussion
- Backpressure, fairness, observability, and safe schedule updates
### Follow-up Questions
- How would you avoid a database scan for every scheduling tick?
- What happens when a worker finishes just after its lease expires?
- How would you roll out a scheduler change without missing or duplicating runs?
- How would you handle daylight-saving transitions for calendar schedules?
Quick Answer: Design a distributed scheduler for one-time and recurring jobs with pause, resume, cancellation, retries, and multi-tenant fairness. Address durable lifecycle state, due-work discovery, worker failure, duplicate delivery, idempotent handlers, backpressure, time-zone semantics, and safe operational rollouts.