I generally don't like judging interviewers or companies, but this company was pretty disgusting. Right from the start, they said that if I didn't interview well, I'd be kicked out early, in the middle of the interview. The person who came in was a technical team lead, looking arrogant. HR was out of their mind and scheduled him for two rounds in a row, one behavioral and one system design. He wasn't happy about it himself and kept acting indifferent, even running out midway to get his laptop power supply.
He barely said anything. He just gave me the question and then went quiet. Most of the content below is what I recalled and filled in afterward.
Distributed Workflow Management System Design Interview Question (Tens of Millions in Scale + Mixed LLM Tasks)
1. Problem Background and Business Scenario
Design a distributed workflow management system that supports tens of millions of workflow definitions and mixes ordinary I/O tasks with large language model (LLM) tasks.
- Task types:
- I/O-intensive tasks: File reads and writes, data cleaning, object-storage uploads and downloads, external database reads and writes, and webhook calls.
- AI / compute-intensive tasks: Chunking long documents, LLM summarization, vector embeddings, entity extraction, and model inference.
- Task dependencies: Tasks are organized in a directed acyclic graph (DAG), with explicit ordering dependencies and contextual data flowing between them.
2. System Scale and Requirements
- Workflow templates: 10 million (10M) predefined workflows.
- Daily execution frequency: About 1 million complete workflows are triggered per day, with an average of 10 subtasks per workflow (8 I/O tasks + 2 LLM tasks), meaning about 100 million task executions scheduled per day.
- Peak scheduling throughput: 5,000–10,000 tasks/second.
- Differences in task execution time:
- Traditional I/O tasks: 10–50 ms.
- LLM batch tasks: 10–60 seconds, depending on input length and model response.
- Data throughput: About 200 TB of raw files and intermediate results per day.
- LLM token throughput: About 80 billion tokens per day (an average of about 925,000 tokens/second), constrained by upstream providers' TPM/RPM quotas.
3. Core Interview Questions
Question 1: Head-of-Line Blocking When Fast and Slow Tasks Run Together
When fast I/O tasks taking milliseconds coexist with slow LLM tasks taking tens of seconds, how would you design task queues and worker scheduling to prevent slow tasks from consuming all worker thread/process slots and starving or severely delaying fast I/O tasks?
Question 2: Storage Design for Massive Metadata and State Logs
With 100 million task state-transition records and logs generated per day (about 200 GB of change records daily), how should the metadata database (RDBMS / NoSQL) design its table schemas, sharding, and storage tiers to support highly concurrent writes and avoid deteriorating query performance over long periods?
Question 3: Strict Upstream LLM API Rate Limits (HTTP 429) and Backpressure
When external LLM services, such as OpenAI and Anthropic, encounter bursts of high concurrency and return HTTP 429 Rate Limit, how should the scheduling engine implement distributed rate limiting, dynamic backoff, and backpressure to prevent cascading failure across the workflow scheduling cluster?
Question 4: Long-Running AI Task Failures and Duplicate Charges (Token Cost Control)
If a long-text summarization task processing hundreds of pages suffers a network disconnection or node failure during its later stages, how would you design fault tolerance and state checkpoints to avoid wasting large amounts of token charges by calling the model again for the entire task on retry?
Question 5: Execution Idempotency Under Network Partitions and Node Failures
When worker crashes, expired leases, or network partitions cause duplicate dispatches, how can the system guarantee end-to-end exactly-once for tasks with irreversible external side effects, such as billing, sending messages, or external writes, under all circumstances at the protocol and storage levels?
Question 6: Very Long Asynchronous Waits and Human Approval (Human-in-the-Loop)
For task nodes that need to wait for an external webhook callback or for human approval days later, such as a compliance review, before execution can continue, how should the system manage resources during the wait to avoid wasting workers?
Question 7: Multi-Tenant Resource Isolation and Noisy Neighbor Prevention
In a multi-tenant environment, if one tenant submits hundreds of thousands of workflow execution requests at once, how should the scheduler implement fair scheduling and quota limits at the queue and worker-node layers to protect other tenants' scheduling SLAs?
Question 8: Unifying Real-Time Typewriter-Style Streaming Output with DAG Batch Delivery
When an interactive frontend needs to display model-generated content in real time through token streaming, but downstream DAG nodes depend on complete structured results, how can the underlying system support both low-latency streaming and reliable batch state commits?
Question 9: Cycle Detection in User-Defined DAGs and Prevention of Runtime Infinite Loops
For complex workflow topologies that users dynamically create or modify, how can the system verify at configuration time, with sub-millisecond latency, that the graph is directed and acyclic? When dynamic conditional branches or retry loops are supported, how can it prevent infinite loops at runtime?
Question 10: State Recovery After a Complete Orchestration and Scheduling Cluster Failure
When every node in the central scheduling engine cluster suddenly crashes or restarts after a power outage, how should the enormous number of running workflow contexts in memory be recovered? How can persistence rebuild the topology of unfinished tasks and resume execution without producing dirty data or duplicate execution?
Discussion
Loading comments…