Design a Scalable Job Scheduler

Quick Overview

This question evaluates system design skills focused on job scheduling, distributed systems concerns (state modeling, retries, idempotency), API and data-model design, fault tolerance, observability, and scalability within the System Design domain.

Design a Scalable Job Scheduler

Company: Nextdoor

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design a job scheduler for a small startup, then explain how you would evolve it to support roughly 100x more scheduled tasks. The system should: - Allow users or internal services to create one-time and recurring jobs. - Run jobs at the scheduled time with reasonable accuracy. - Support retries for failed jobs. - Support cancellation and rescheduling. - Track job states such as pending, dispatched, running, succeeded, and failed. - Provide basic observability for operators. Discuss: 1. API design and core data model. 2. How jobs are stored and selected for execution. 3. How workers execute jobs safely. 4. How to handle duplicate execution, retries, and idempotency. 5. Failure handling when the scheduler node or worker crashes. 6. How the initial design for a low-scale startup would work. 7. What architectural changes you would make to scale to 100x more jobs.

Overview: This question evaluates system design skills focused on job scheduling, distributed systems concerns (state modeling, retries, idempotency), API and data-model design, fault tolerance, observability, and scalability within the System Design domain.

Community answers

Answer by lionel.eisenberg

Transitions — who writes each edge, and what makes it safe The state names are worth nothing. The transitions are the whole answer. | Edge | Writer | Safety mechanism | | --- | --- | --- | | ∅ → PENDING | API | UNIQUE (tenant_id, idempotency_key) and UNIQUE (job_id, scheduled_for). The index is the arbiter, not an application-level check. | | PENDING → CLAIMED | Dispatcher / worker | Conditional update WHERE state='PENDING' AND run_at <= now() under FOR UPDATE SKIP LOCKED. Stamps lease_owner, a fresh lease_token, lease_expires_at, and increments attempt. | | CLAIMED → RUNNING | Worker | Conditional on lease_token matching and lease_expires_at > now(). | | RUNNING → SUCCEEDED / RETRY_WAIT | Worker | Fenced ack. WHERE run_id=? AND lease_token=? AND lease_expires_at > now(). If it updates zero rows, the worker lost the lease and must not assume its work counted. | | CLAIMED / RUNNING → PENDING | Reaper | WHERE lease_expires_at < now(). This is the only edge that can cause duplicate execution — and it's deliberate. Its rate is a metric you alert on. |
|Home/System Design/Nextdoor
Nextdoor logo
Nextdoor
Mar 16, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
28
0

Design a job scheduler for a small startup, then explain how you would evolve it to support roughly 100x more scheduled tasks.

The system should:

  • Allow users or internal services to create one-time and recurring jobs.
  • Run jobs at the scheduled time with reasonable accuracy.
  • Support retries for failed jobs.
  • Support cancellation and rescheduling.
  • Track job states such as pending, dispatched, running, succeeded, and failed.
  • Provide basic observability for operators.

Discuss:

  1. API design and core data model.
  2. How jobs are stored and selected for execution.
  3. How workers execute jobs safely.
  4. How to handle duplicate execution, retries, and idempotency.
  5. Failure handling when the scheduler node or worker crashes.
  6. How the initial design for a low-scale startup would work.
  7. What architectural changes you would make to scale to 100x more jobs.

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...