Design a Reliable Job Scheduler
Company: Nuro
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design a general-purpose job scheduler. Clients submit jobs for execution at or after a requested time, inspect status, and cancel jobs that have not completed. Workers execute jobs and report outcomes.
### Constraints & Assumptions
- Delivery may be at least once; a worker can crash after performing work but before acknowledging it.
- Jobs may be scheduled far in the future.
- A job may have a runtime timeout and a retry policy.
- Exact scale, scheduling precision, and dependency support are not prescribed; ask before choosing storage or partitioning.
### Clarifying Questions to Ask
- Is this one-time scheduling, recurring scheduling, dependency scheduling, or all three?
- What timing precision and lateness are acceptable?
- Must jobs run in submission order or by priority?
- Can handlers be made idempotent?
- What cancellation behavior is expected for a running job?
### What a Strong Answer Covers
- Durable job state and explicit transitions
- Efficient discovery of due jobs
- Leases, retries, and worker crash recovery
- Idempotency and duplicate-execution limits
- Fairness, backpressure, and observability
- Cancellation races and dead-letter handling
### Follow-up Questions
- Add recurring schedules without double-firing.
- Support dependencies between jobs.
- Prevent one tenant from monopolizing workers.
Quick Answer: Design a reliable job scheduler for delayed execution, status inspection, cancellation, timeouts, and retries across fallible workers. Clarify schedule types and precision, then reason about duplicate execution, crash recovery, cancellation races, fairness, backpressure, dependencies, tenant isolation, and operational visibility.