Debug a Python Job Scheduler: Failed Jobs Stuck in PENDING and Async Jobs Never Run

Read the full interview experience this question came from →

Quick Overview

Debug an existing Python job scheduler from bug reports and write a regression test for every fix. A threaded, heap-based scheduler leaves failing jobs stuck in PENDING, and an asyncio-based scheduler seems never to run submitted jobs; tests root-cause analysis, failure handling and deterministic concurrent tests.

Debug a Python Job Scheduler: Failed Jobs Stuck in PENDING and Async Jobs Never Run

Company: OpenAI

Role: Software Engineer

Category: Software Engineering Fundamentals

Difficulty: hard

Interview Round: Onsite

You are given an existing Python job scheduler repository and a set of bug reports. For each report, read the code, find the cause, fix it, and write a test for the fix yourself. Web search is allowed. The exercise is a debugging session on an unfamiliar codebase rather than writing code from scratch. ### Constraints and Clarifications - Every fix must come with a test that you write. - The original repository is not reproduced here. The snippet in Part 1 is a minimal reconstruction of the described behavior. The code for Part 2 is not available, so that part is a diagnosis exercise. ### Clarifying Questions - How do callers observe a job's progress: by polling `status`, by waiting on the job, or through callbacks? - Is there an existing test suite and test runner to extend? - Should fixes be minimal, or may you restructure code that the bug report does not mention? ### Part 1 — In-memory scheduler: failed jobs stay PENDING A `Job` holds a function, its `args` and `kwargs`, an execution time `run_at`, and a `status`. The scheduler keeps jobs in a heap ordered by `run_at`, takes out the jobs that are due, and runs each one on its own thread. Bug report: a job whose function returns normally is marked `COMPLETED`, but when the function raises an exception, the status update after the call never runs and the job stays `PENDING` forever. ```python import heapq import itertools import threading import time from dataclasses import dataclass from enum import Enum from typing import Any, Callable class Status(Enum): PENDING = "PENDING" COMPLETED = "COMPLETED" @dataclass class Job: func: Callable[..., Any] args: tuple kwargs: dict run_at: float # seconds since the epoch, as returned by time.time() status: Status = Status.PENDING class Scheduler: def __init__(self) -> None: self._heap: list[tuple[float, int, Job]] = [] self._seq = itertools.count() # tie-breaker for equal run_at values self._lock = threading.Lock() def schedule(self, func: Callable[..., Any], run_at: float, *args: Any, **kwargs: Any) -> Job: job = Job(func, args, kwargs, run_at) with self._lock: heapq.heappush(self._heap, (run_at, next(self._seq), job)) return job def run_pending(self) -> None: """Start every job whose run_at has passed, each on its own thread.""" now = time.time() while True: with self._lock: if not self._heap or self._heap[0][0] > now: return _, _, job = heapq.heappop(self._heap) threading.Thread(target=self._execute, args=(job,)).start() def _execute(self, job: Job) -> None: job.func(*job.args, **job.kwargs) job.status = Status.COMPLETED ``` Find the root cause, write a test that fails because of it, and fix it. ```hint Follow the exception Trace what happens to an exception raised inside the job's function: which statements after the call still run, and where the exception ends up when it leaves a worker thread. ``` ```hint A test that does not sleep A test that sleeps and hopes the thread has finished is flaky. Find a way to exercise the failing path deterministically, or to know for certain that the job has finished before you assert. ``` #### Clarifying Questions for this Part - Is there a status for failed jobs, or may you add one? Which other code reads job statuses and would need to handle it? - Should the exception be recorded on the job, logged, re-raised, or retried? #### What This Part Should Cover - The root cause: why the status assignment is skipped, and why nobody notices the failure - A fix that leaves every job in a terminal status on every path and keeps the error - A regression test that fails before the fix and passes after it, without timing-based flakiness - Whether the fix changes behavior for jobs that succeed ### Part 2 — Async scheduler: submitted jobs never seem to run The repository also contains an asyncio-based scheduler. Its bug report says that jobs submitted to it appear never to actually run. Explain how you would reproduce and narrow down this report, which root causes in an asyncio scheduler can produce this symptom and how you would tell them apart, how you would fix the cause you find, and what test would prove the fix. ```hint Created is not running In asyncio, calling a coroutine function or creating a task does not by itself guarantee that the job's body ever executes. List everything that must be true for a submitted job's body to start. ``` ```hint Let the runtime help asyncio has a debug mode, and Python emits specific warnings for some of these mistakes. Think about which observations would tell the candidate causes apart. ``` #### Clarifying Questions for this Part - Are jobs submitted from code running inside the event loop, or from other threads? - Are job functions coroutine functions, plain functions, or both? - Does "never run" mean no side effects at all, or that the status never changes? #### What This Part Should Cover - A minimal reproduction and a systematic way to narrow down the cause - Concrete asyncio failure modes that leave submitted work unexecuted, with the signal that identifies each - A fix for the cause found, with its regression test - How failures in background tasks become visible instead of silent ### What a Strong Answer Covers - A failing test written for each bug report before the code changes - Root causes explained precisely rather than patched by symptom - Every job ends in a terminal status, including on failure paths - Deterministic tests for threaded and asynchronous code - Small fixes that do not change unrelated behavior ### Follow-up Questions - A job's function hangs forever. How would each scheduler detect it, and what can it actually do about the running work? - Should failed jobs be retried? Where would the retry policy live, and which jobs must never be retried automatically? - The threaded scheduler starts one thread per due job. What happens when thousands of jobs become due at once, and what would you change?

Overview: Debug an existing Python job scheduler from bug reports and write a regression test for every fix. A threaded, heap-based scheduler leaves failing jobs stuck in PENDING, and an asyncio-based scheduler seems never to run submitted jobs; tests root-cause analysis, failure handling and deterministic concurrent tests.

Read the full OpenAI Software Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/OpenAI
OpenAI logo
OpenAI
Jul 29, 2026
hardSoftware EngineerOnsiteSoftware Engineering Fundamentals
0
0

You are given an existing Python job scheduler repository and a set of bug reports. For each report, read the code, find the cause, fix it, and write a test for the fix yourself. Web search is allowed. The exercise is a debugging session on an unfamiliar codebase rather than writing code from scratch.

Constraints and Clarifications

  • Every fix must come with a test that you write.
  • The original repository is not reproduced here. The snippet in Part 1 is a minimal reconstruction of the described behavior. The code for Part 2 is not available, so that part is a diagnosis exercise.

Clarifying Questions Guidance

  • How do callers observe a job's progress: by polling status , by waiting on the job, or through callbacks?
  • Is there an existing test suite and test runner to extend?
  • Should fixes be minimal, or may you restructure code that the bug report does not mention?

Part 1 — In-memory scheduler: failed jobs stay PENDING

A Job holds a function, its args and kwargs, an execution time run_at, and a status. The scheduler keeps jobs in a heap ordered by run_at, takes out the jobs that are due, and runs each one on its own thread.

Bug report: a job whose function returns normally is marked COMPLETED, but when the function raises an exception, the status update after the call never runs and the job stays PENDING forever.

import heapq
import itertools
import threading
import time
from dataclasses import dataclass
from enum import Enum
from typing import Any, Callable


class Status(Enum):
    PENDING = "PENDING"
    COMPLETED = "COMPLETED"


@dataclass
class Job:
    func: Callable[..., Any]
    args: tuple
    kwargs: dict
    run_at: float                    # seconds since the epoch, as returned by time.time()
    status: Status = Status.PENDING


class Scheduler:
    def __init__(self) -> None:
        self._heap: list[tuple[float, int, Job]] = []
        self._seq = itertools.count()          # tie-breaker for equal run_at values
        self._lock = threading.Lock()

    def schedule(self, func: Callable[..., Any], run_at: float, *args: Any, **kwargs: Any) -> Job:
        job = Job(func, args, kwargs, run_at)
        with self._lock:
            heapq.heappush(self._heap, (run_at, next(self._seq), job))
        return job

    def run_pending(self) -> None:
        """Start every job whose run_at has passed, each on its own thread."""
        now = time.time()
        while True:
            with self._lock:
                if not self._heap or self._heap[0][0] > now:
                    return
                _, _, job = heapq.heappop(self._heap)
            threading.Thread(target=self._execute, args=(job,)).start()

    def _execute(self, job: Job) -> None:
        job.func(*job.args, **job.kwargs)
        job.status = Status.COMPLETED

Find the root cause, write a test that fails because of it, and fix it.

Clarifying Questions for this Part Guidance

  • Is there a status for failed jobs, or may you add one? Which other code reads job statuses and would need to handle it?
  • Should the exception be recorded on the job, logged, re-raised, or retried?

What This Part Should Cover Guidance

  • The root cause: why the status assignment is skipped, and why nobody notices the failure
  • A fix that leaves every job in a terminal status on every path and keeps the error
  • A regression test that fails before the fix and passes after it, without timing-based flakiness
  • Whether the fix changes behavior for jobs that succeed

Part 2 — Async scheduler: submitted jobs never seem to run

The repository also contains an asyncio-based scheduler. Its bug report says that jobs submitted to it appear never to actually run. Explain how you would reproduce and narrow down this report, which root causes in an asyncio scheduler can produce this symptom and how you would tell them apart, how you would fix the cause you find, and what test would prove the fix.

Clarifying Questions for this Part Guidance

  • Are jobs submitted from code running inside the event loop, or from other threads?
  • Are job functions coroutine functions, plain functions, or both?
  • Does "never run" mean no side effects at all, or that the status never changes?

What This Part Should Cover Guidance

  • A minimal reproduction and a systematic way to narrow down the cause
  • Concrete asyncio failure modes that leave submitted work unexecuted, with the signal that identifies each
  • A fix for the cause found, with its regression test
  • How failures in background tasks become visible instead of silent

What a Strong Answer Covers Guidance

  • A failing test written for each bug report before the code changes
  • Root causes explained precisely rather than patched by symptom
  • Every job ends in a terminal status, including on failure paths
  • Deterministic tests for threaded and asynchronous code
  • Small fixes that do not change unrelated behavior

Follow-up Questions Guidance

  • A job's function hangs forever. How would each scheduler detect it, and what can it actually do about the running work?
  • Should failed jobs be retried? Where would the retry policy live, and which jobs must never be retried automatically?
  • The threaded scheduler starts one thread per due job. What happens when thousands of jobs become due at once, and what would you change?
Loading comments...