Recover DAG workflows through crashes, duplicate attempts, external side effects, and multi-day approvals using durable state and explicit exactly-once limits.
Recover Durable Workflows with External Effects and Long Waits
Company: Qualified Health
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
Design durable execution and recovery for a distributed DAG workflow system. Workers can crash or lose connectivity, tasks can have irreversible external effects such as billing or sending a message, and some nodes wait for a webhook or human approval for several days. The entire scheduling cluster may restart while many runs are unfinished.
### Constraints & Assumptions
- A lease expiry can cause duplicate dispatch even while an old worker continues running.
- The source asks for end-to-end exactly-once external effects. Evaluate whether this is achievable for each external service contract rather than assuming the scheduler alone can guarantee it.
- Waiting for a callback or approval must not occupy a worker thread for days.
- Recovery must preserve the run's dependency graph, completed outputs, pending work, and unresolved external operations.
### Clarifying Questions to Ask
- Can each external destination deduplicate operations or participate in a suitable transaction?
- Can the status of a timed-out side effect be queried by a stable operation ID?
- Are workflow definitions immutable for an existing run, and where are task outputs persisted?
- How are callbacks authenticated, correlated, deduplicated, and ordered against timeouts or cancellation?
### Part 1 — External Effects and Duplicate Attempts
Describe durable intent, attempt ownership, completion, and retry behavior. Explain the difference between exactly-once logical state transitions and exactly-once effects at an external destination.
#### What This Part Should Cover
- Stable logical operation identity separate from worker attempt identity.
- Fenced local writes and destination-enforced deduplication where available.
- An explicit uncertain state when neither safe retry nor confirmed success is established.
### Part 2 — Durable Waits
Design webhook and human-approval nodes that can wait without retaining an executing worker. Explain registration, callback races, repeated callbacks, deadlines, and resumption.
#### What This Part Should Cover
- Persisted correlation, authorization, deadline, and wait state.
- Atomic selection of one terminal wait outcome.
- Durable readiness of downstream tasks after resumption.
### Part 3 — Full Scheduler Recovery
Explain how a replacement scheduling cluster reconstructs unfinished runs after complete loss of scheduler memory. Distinguish tasks that can be dispatched safely from attempts whose external outcomes remain uncertain.
#### What This Part Should Cover
- Definition versions, task states, dependency completion, and durable artifacts.
- Recovery ownership and bounded scans or replay.
- No blind reset of every running task to a fresh execution.
```hint Separate a job from an attempt
The logical external operation should retain its identity even if the scheduler assigns a replacement worker. A new worker attempt is not automatically a new billable or message-sending operation.
```
### What a Strong Answer Covers
- Durable state that survives worker and scheduler failure.
- Honest exactly-once boundaries and safe treatment of uncertain external outcomes.
- Resource-free long waits with race-safe resumption.
- Recovery that rebuilds readiness without losing or duplicating logical completion.
### Follow-up Questions
- What if approval arrives at the same time as its deadline expires?
- Why is a transactional outbox insufficient by itself to deduplicate an external service that ignores operation IDs?
- How would you prevent two recovery schedulers from resuming the same run independently?
Overview: Recover DAG workflows through crashes, duplicate attempts, external side effects, and multi-day approvals using durable state and explicit exactly-once limits.
Recover Durable Workflows with External Effects and Long Waits
Qualified Health
Sep 30, 2026
hardSoftware EngineerOnsiteSystem Design
0
0
Design durable execution and recovery for a distributed DAG workflow system. Workers can crash or lose connectivity, tasks can have irreversible external effects such as billing or sending a message, and some nodes wait for a webhook or human approval for several days. The entire scheduling cluster may restart while many runs are unfinished.
Constraints & Assumptions
A lease expiry can cause duplicate dispatch even while an old worker continues running.
The source asks for end-to-end exactly-once external effects. Evaluate whether this is achievable for each external service contract rather than assuming the scheduler alone can guarantee it.
Waiting for a callback or approval must not occupy a worker thread for days.
Recovery must preserve the run's dependency graph, completed outputs, pending work, and unresolved external operations.
Clarifying Questions to Ask Guidance
Can each external destination deduplicate operations or participate in a suitable transaction?
Can the status of a timed-out side effect be queried by a stable operation ID?
Are workflow definitions immutable for an existing run, and where are task outputs persisted?
How are callbacks authenticated, correlated, deduplicated, and ordered against timeouts or cancellation?
Part 1 — External Effects and Duplicate Attempts
Describe durable intent, attempt ownership, completion, and retry behavior. Explain the difference between exactly-once logical state transitions and exactly-once effects at an external destination.
What This Part Should Cover Guidance
Stable logical operation identity separate from worker attempt identity.
Fenced local writes and destination-enforced deduplication where available.
An explicit uncertain state when neither safe retry nor confirmed success is established.
Part 2 — Durable Waits
Design webhook and human-approval nodes that can wait without retaining an executing worker. Explain registration, callback races, repeated callbacks, deadlines, and resumption.
What This Part Should Cover Guidance
Persisted correlation, authorization, deadline, and wait state.
Atomic selection of one terminal wait outcome.
Durable readiness of downstream tasks after resumption.
Part 3 — Full Scheduler Recovery
Explain how a replacement scheduling cluster reconstructs unfinished runs after complete loss of scheduler memory. Distinguish tasks that can be dispatched safely from attempts whose external outcomes remain uncertain.
What This Part Should Cover Guidance
Definition versions, task states, dependency completion, and durable artifacts.
Recovery ownership and bounded scans or replay.
No blind reset of every running task to a fresh execution.
What a Strong Answer Covers Guidance
Durable state that survives worker and scheduler failure.
Honest exactly-once boundaries and safe treatment of uncertain external outcomes.
Resource-free long waits with race-safe resumption.
Recovery that rebuilds readiness without losing or duplicating logical completion.
Follow-up Questions Guidance
What if approval arrives at the same time as its deadline expires?
Why is a transactional outbox insufficient by itself to deduplicate an external service that ignores operation IDs?
How would you prevent two recovery schedulers from resuming the same run independently?