Design a workflow orchestration system that runs multi-step jobs reliably

Read the full interview experience this question came from →

Quick Overview

A system design question asking you to build a workflow system where users define multi-step workflows with dependencies, start runs, and track progress. It tests data modeling, orchestration and worker architecture, durable state, retries and idempotency, fairness between tenants, scaling and observability.

Design a workflow orchestration system that runs multi-step jobs reliably

Company: Mercor

Role: Machine Learning Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a workflow system. Users define a workflow as a set of steps with dependencies between them, start runs of that workflow, and the system executes each step in the right order, tracks the state of every run and step, handles failures, and lets users see progress and results. ```hint Crash in the middle Ask what must survive if the machine coordinating a run dies halfway through a step. ``` ```hint Deciding versus doing Consider which component decides what runs next and which component actually runs a step. ``` ### Constraints and Clarifications - No scale figures, step types or trigger mechanisms were given; establish them before designing. - Assume steps are automated tasks (calls to services, scripts or containerized jobs) unless the interviewer puts human steps in scope. ### Clarifying Questions - Are steps automated jobs, human tasks such as reviews or approvals, or both? - How are runs started: an API call, a schedule, or external events? - Roughly how many workflow definitions, concurrent runs and steps per run are expected, and how long do steps take (seconds, hours, days)? - What failure semantics are required: retries, timeouts, skipping, compensation, manual intervention? - Are workflows linear, or do they need branching, fan-out and fan-in? - Is the system shared by many teams, and must they be isolated from each other's load? ### What a Strong Answer Covers - Scoped functional and non-functional requirements - A data model for versioned definitions, runs and step executions, with an explicit state machine - Architecture: API, orchestration logic, durable task queue, workers and state store, and how a step moves from ready to complete - Reliability: durable state, at-least-once execution with idempotent steps, retries with backoff, timeouts and heartbeats, and recovery after crashes - Scaling, fairness between tenants, and observability ### Follow-up Questions - A worker completes a step's side effect and then crashes before reporting success. How do you avoid doing the side effect twice? - How do users change a workflow definition while runs of the old version are still in flight? - One tenant starts a massive fan-out and starves everyone else. What controls do you add? - How would you support a step that waits days for an external signal such as a human approval?

Overview: A system design question asking you to build a workflow system where users define multi-step workflows with dependencies, start runs, and track progress. It tests data modeling, orchestration and worker architecture, durable state, retries and idempotency, fairness between tenants, scaling and observability.

Read the full Mercor Machine Learning Engineer interview experience this question came from

|Home/System Design/Mercor
Mercor logo
Mercor
Sep 25, 2026
mediumMachine Learning EngineerOnsiteSystem Design
1
0

Design a workflow system. Users define a workflow as a set of steps with dependencies between them, start runs of that workflow, and the system executes each step in the right order, tracks the state of every run and step, handles failures, and lets users see progress and results.

Constraints and Clarifications

  • No scale figures, step types or trigger mechanisms were given; establish them before designing.
  • Assume steps are automated tasks (calls to services, scripts or containerized jobs) unless the interviewer puts human steps in scope.

Clarifying Questions Guidance

  • Are steps automated jobs, human tasks such as reviews or approvals, or both?
  • How are runs started: an API call, a schedule, or external events?
  • Roughly how many workflow definitions, concurrent runs and steps per run are expected, and how long do steps take (seconds, hours, days)?
  • What failure semantics are required: retries, timeouts, skipping, compensation, manual intervention?
  • Are workflows linear, or do they need branching, fan-out and fan-in?
  • Is the system shared by many teams, and must they be isolated from each other's load?

What a Strong Answer Covers Guidance

  • Scoped functional and non-functional requirements
  • A data model for versioned definitions, runs and step executions, with an explicit state machine
  • Architecture: API, orchestration logic, durable task queue, workers and state store, and how a step moves from ready to complete
  • Reliability: durable state, at-least-once execution with idempotent steps, retries with backoff, timeouts and heartbeats, and recovery after crashes
  • Scaling, fairness between tenants, and observability

Follow-up Questions Guidance

  • A worker completes a step's side effect and then crashes before reporting success. How do you avoid doing the side effect twice?
  • How do users change a workflow definition while runs of the old version are still in flight?
  • One tenant starts a massive fan-out and starves everyone else. What controls do you add?
  • How would you support a step that waits days for an external signal such as a human approval?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...