Implement a Lightweight ETL Pipeline Scheduler in Python

Read the full interview experience this question came from →

Quick Overview

Implement a small Python scheduler that registers ETL tasks with dependencies and runs a pipeline so that each task executes only after its upstream tasks succeed. It tests clean interface design, dependency ordering with validation, failure and retry handling, and clear reporting of what happened to each task.

Implement a Lightweight ETL Pipeline Scheduler in Python

Company: Unity

Role: Software Engineer

Category: Software Engineering Fundamentals

Difficulty: medium

Interview Round: Onsite

Implement, in Python, a lightweight scheduler for an ETL (extract, transform, load) pipeline. The interviewer described the logic as simple; the round is about producing clean, working Python within the interview. The task was not specified further, so confirm this baseline with the interviewer before coding: - A pipeline is a set of named tasks. Each task is a Python callable, such as an extract, a transform or a load step, and a task may depend on other tasks. - The scheduler lets a caller register tasks together with their dependencies, and then run the pipeline once. - A task runs only after every task it depends on has finished successfully, and each task runs at most once per run. - At the end of a run, the caller can see what happened to every task. ```hint Who can go first Think about which tasks are ready at the very start, and how finishing one task changes which others become ready. ``` ```hint Decide what failure means Before coding, decide what happens to the tasks downstream of a task that raises an exception, and what the caller sees at the end. ``` ### Constraints and Clarifications - Use only the Python standard library. - A pipeline whose dependencies cannot all be satisfied should be detected rather than left to hang or loop. ### Clarifying Questions - Should the scheduler also trigger runs on a time schedule (for example, every hour), or only run the pipeline when called? - When a task fails, should it be retried, and should the rest of the pipeline stop or continue with tasks that do not depend on it? - Should independent tasks run in parallel, and if so with what concurrency limit? - Do tasks pass data to each other, or do they communicate only through external storage? - What should happen if the dependencies form a cycle or name a task that does not exist? ### What a Strong Answer Covers - A clear task and scheduler interface, with the baseline confirmed before coding - Correct dependency ordering, including validation of cycles and unknown dependencies - Defined failure behavior: exceptions captured, downstream tasks skipped, optional retries - A run report that tells the caller what happened to every task - Readable, idiomatic Python with a small usage example and a complexity statement ### Follow-up Questions - Run independent tasks in parallel with at most N running at a time. What changes in your design? - Add retries with exponential backoff. Which kinds of failure should not be retried? - Make the pipeline run every hour. What happens if a run takes longer than an hour? - The process crashes halfway through a run. How do you resume without loading the same data twice?

Overview: Implement a small Python scheduler that registers ETL tasks with dependencies and runs a pipeline so that each task executes only after its upstream tasks succeed. It tests clean interface design, dependency ordering with validation, failure and retry handling, and clear reporting of what happened to each task.

Read the full Unity Software Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/Unity
Unity logo
Unity
Oct 9, 2026
mediumSoftware EngineerOnsiteSoftware Engineering Fundamentals
1
0

Implement, in Python, a lightweight scheduler for an ETL (extract, transform, load) pipeline. The interviewer described the logic as simple; the round is about producing clean, working Python within the interview.

The task was not specified further, so confirm this baseline with the interviewer before coding:

  • A pipeline is a set of named tasks. Each task is a Python callable, such as an extract, a transform or a load step, and a task may depend on other tasks.
  • The scheduler lets a caller register tasks together with their dependencies, and then run the pipeline once.
  • A task runs only after every task it depends on has finished successfully, and each task runs at most once per run.
  • At the end of a run, the caller can see what happened to every task.

Constraints and Clarifications

  • Use only the Python standard library.
  • A pipeline whose dependencies cannot all be satisfied should be detected rather than left to hang or loop.

Clarifying Questions Guidance

  • Should the scheduler also trigger runs on a time schedule (for example, every hour), or only run the pipeline when called?
  • When a task fails, should it be retried, and should the rest of the pipeline stop or continue with tasks that do not depend on it?
  • Should independent tasks run in parallel, and if so with what concurrency limit?
  • Do tasks pass data to each other, or do they communicate only through external storage?
  • What should happen if the dependencies form a cycle or name a task that does not exist?

What a Strong Answer Covers Guidance

  • A clear task and scheduler interface, with the baseline confirmed before coding
  • Correct dependency ordering, including validation of cycles and unknown dependencies
  • Defined failure behavior: exceptions captured, downstream tasks skipped, optional retries
  • A run report that tells the caller what happened to every task
  • Readable, idiomatic Python with a small usage example and a complexity statement

Follow-up Questions Guidance

  • Run independent tasks in parallel with at most N running at a time. What changes in your design?
  • Add retries with exponential backoff. Which kinds of failure should not be retried?
  • Make the pipeline run every hour. What happens if a run takes longer than an hour?
  • The process crashes halfway through a run. How do you resume without loading the same data twice?
Loading comments...