Debug a Python Job Scheduler: Failed Jobs Stuck in PENDING and Async Jobs Never Run
Company: OpenAI
Role: Software Engineer
Category: Software Engineering Fundamentals
Difficulty: hard
Interview Round: Onsite
You are given an existing Python job scheduler repository and a set of bug reports. For each report, read the code, find the cause, fix it, and write a test for the fix yourself. Web search is allowed. The exercise is a debugging session on an unfamiliar codebase rather than writing code from scratch.
### Constraints and Clarifications
- Every fix must come with a test that you write.
- The original repository is not reproduced here. The snippet in Part 1 is a minimal reconstruction of the described behavior. The code for Part 2 is not available, so that part is a diagnosis exercise.
### Clarifying Questions
- How do callers observe a job's progress: by polling `status`, by waiting on the job, or through callbacks?
- Is there an existing test suite and test runner to extend?
- Should fixes be minimal, or may you restructure code that the bug report does not mention?
### Part 1 — In-memory scheduler: failed jobs stay PENDING
A `Job` holds a function, its `args` and `kwargs`, an execution time `run_at`, and a `status`. The scheduler keeps jobs in a heap ordered by `run_at`, takes out the jobs that are due, and runs each one on its own thread.
Bug report: a job whose function returns normally is marked `COMPLETED`, but when the function raises an exception, the status update after the call never runs and the job stays `PENDING` forever.
```python
import heapq
import itertools
import threading
import time
from dataclasses import dataclass
from enum import Enum
from typing import Any, Callable
class Status(Enum):
PENDING = "PENDING"
COMPLETED = "COMPLETED"
@dataclass
class Job:
func: Callable[..., Any]
args: tuple
kwargs: dict
run_at: float # seconds since the epoch, as returned by time.time()
status: Status = Status.PENDING
class Scheduler:
def __init__(self) -> None:
self._heap: list[tuple[float, int, Job]] = []
self._seq = itertools.count() # tie-breaker for equal run_at values
self._lock = threading.Lock()
def schedule(self, func: Callable[..., Any], run_at: float, *args: Any, **kwargs: Any) -> Job:
job = Job(func, args, kwargs, run_at)
with self._lock:
heapq.heappush(self._heap, (run_at, next(self._seq), job))
return job
def run_pending(self) -> None:
"""Start every job whose run_at has passed, each on its own thread."""
now = time.time()
while True:
with self._lock:
if not self._heap or self._heap[0][0] > now:
return
_, _, job = heapq.heappop(self._heap)
threading.Thread(target=self._execute, args=(job,)).start()
def _execute(self, job: Job) -> None:
job.func(*job.args, **job.kwargs)
job.status = Status.COMPLETED
```
Find the root cause, write a test that fails because of it, and fix it.
```hint Follow the exception
Trace what happens to an exception raised inside the job's function: which statements after the call still run, and where the exception ends up when it leaves a worker thread.
```
```hint A test that does not sleep
A test that sleeps and hopes the thread has finished is flaky. Find a way to exercise the failing path deterministically, or to know for certain that the job has finished before you assert.
```
#### Clarifying Questions for this Part
- Is there a status for failed jobs, or may you add one? Which other code reads job statuses and would need to handle it?
- Should the exception be recorded on the job, logged, re-raised, or retried?
#### What This Part Should Cover
- The root cause: why the status assignment is skipped, and why nobody notices the failure
- A fix that leaves every job in a terminal status on every path and keeps the error
- A regression test that fails before the fix and passes after it, without timing-based flakiness
- Whether the fix changes behavior for jobs that succeed
### Part 2 — Async scheduler: submitted jobs never seem to run
The repository also contains an asyncio-based scheduler. Its bug report says that jobs submitted to it appear never to actually run. Explain how you would reproduce and narrow down this report, which root causes in an asyncio scheduler can produce this symptom and how you would tell them apart, how you would fix the cause you find, and what test would prove the fix.
```hint Created is not running
In asyncio, calling a coroutine function or creating a task does not by itself guarantee that the job's body ever executes. List everything that must be true for a submitted job's body to start.
```
```hint Let the runtime help
asyncio has a debug mode, and Python emits specific warnings for some of these mistakes. Think about which observations would tell the candidate causes apart.
```
#### Clarifying Questions for this Part
- Are jobs submitted from code running inside the event loop, or from other threads?
- Are job functions coroutine functions, plain functions, or both?
- Does "never run" mean no side effects at all, or that the status never changes?
#### What This Part Should Cover
- A minimal reproduction and a systematic way to narrow down the cause
- Concrete asyncio failure modes that leave submitted work unexecuted, with the signal that identifies each
- A fix for the cause found, with its regression test
- How failures in background tasks become visible instead of silent
### What a Strong Answer Covers
- A failing test written for each bug report before the code changes
- Root causes explained precisely rather than patched by symptom
- Every job ends in a terminal status, including on failure paths
- Deterministic tests for threaded and asynchronous code
- Small fixes that do not change unrelated behavior
### Follow-up Questions
- A job's function hangs forever. How would each scheduler detect it, and what can it actually do about the running work?
- Should failed jobs be retried? Where would the retry policy live, and which jobs must never be retried automatically?
- The threaded scheduler starts one thread per due job. What happens when thousands of jobs become due at once, and what would you change?
Overview: Debug an existing Python job scheduler from bug reports and write a regression test for every fix. A threaded, heap-based scheduler leaves failing jobs stuck in PENDING, and an asyncio-based scheduler seems never to run submitted jobs; tests root-cause analysis, failure handling and deterministic concurrent tests.
Read the full OpenAI Software Engineer interview experience this question came from