Scale Mixed I/O and LLM Workflow Execution

Read the full interview experience this question came from →

Quick Overview

Scale mixed I/O and LLM workflows with separate worker budgets, tenant fairness, tiered state storage, and explicit reconciliation of conflicting workload figures.

Scale Mixed I/O and LLM Workflow Execution

Company: Qualified Health

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

Design the capacity, queueing, worker isolation, and storage layers of a multi-tenant workflow system whose DAGs mix fast I/O tasks with much slower LLM tasks. ### Reported Workload and Constraints | Item | Reported figure | |---|---| | Stored workflow definitions | 10 million | | Workflow runs | About 1 million per day | | Tasks per run | Average 10: 8 I/O tasks and 2 LLM tasks | | Separately stated daily task volume | About 100 million task executions | | Peak dispatch | 5,000–10,000 tasks per second | | I/O task duration | 10–50 milliseconds | | LLM task duration | 10–60 seconds | | Files and intermediate data | About 200 TB per day | | State changes and logs | About 100 million records, or 200 GB, per day | The run count multiplied by the average task count implies **10 million**, not 100 million, daily task executions. Preserve this discrepancy as a sizing question. A state-transition record also need not correspond one-to-one with a task execution. These figures are interview inputs, not verified measurements of a deployed system. ### Clarifying Questions to Ask - Which daily execution count is authoritative, and how many state events does one attempt generate? - Do task classes have the same proportions at peak as across the day? - What retention and query patterns apply to definitions, live run state, task history, logs, and large intermediate objects? - Which fairness or latency guarantees apply per tenant, and can unused capacity be borrowed? ### Part 1 — Size and Separate Fast and Slow Work Estimate average dispatch rates under both stated daily volumes and explain how runtime differences affect in-flight concurrency. Design queues and worker pools that prevent slow LLM calls from exhausting the slots needed by fast I/O tasks. #### What This Part Should Cover - Correct arithmetic and explicit uncertainty rather than silently choosing a workload figure. - Separate admission, concurrency, and backpressure for task classes. - Capacity estimates that distinguish averages, bursts, and provider-limited throughput. ### Part 2 — Store Definitions, State, Logs, and Payloads Describe records, indexes, partition keys, retention, and storage tiers for the reported scale. Explain how to query a run's current progress without scanning its entire historical log or carrying large document bodies in the metadata database. #### What This Part Should Cover - Immutable definition versions and durable run/task/attempt identities. - Hot state versus append-only history and large object payloads. - Partitioning, bounded indexes, retention, and recovery-safe archival. ### Part 3 — Protect Tenants from Bursty Neighbors A tenant submits hundreds of thousands of runs at once. Explain admission and scheduling controls that protect other tenants while allowing useful work to continue. #### What This Part Should Cover - Tenant limits on submissions, queued work, running tasks, and scarce downstream resources. - Fair dispatch within each task class and controlled borrowing of spare capacity. - Metrics that expose class-specific and tenant-specific starvation. ```hint Compare slots, not just task counts A task that occupies a slot for tens of seconds has a different concurrency cost from one that finishes in milliseconds, even if both count as one dispatched task. ``` ### What a Strong Answer Covers - Capacity reasoning that keeps the conflicting daily figures visible. - Isolation of fast I/O work, slow model calls, tenant demand, and downstream bottlenecks. - Storage and retention choices that preserve current-state reads and recovery without putting daily payload traffic in the scheduling database. ### Follow-up Questions - How would a hot tenant change the effectiveness of sharding only by tenant ID? - What happens to admission when object storage or the LLM provider is slower than the workers? - Which state must remain available before historical events can be archived or removed?

Overview: Scale mixed I/O and LLM workflows with separate worker budgets, tenant fairness, tiered state storage, and explicit reconciliation of conflicting workload figures.

Read the full Qualified Health Software Engineer interview experience this question came from

|Home/System Design/Qualified Health
Qualified Health logo
Qualified Health
Sep 30, 2026
hardSoftware EngineerOnsiteSystem Design
0
0

Design the capacity, queueing, worker isolation, and storage layers of a multi-tenant workflow system whose DAGs mix fast I/O tasks with much slower LLM tasks.

Reported Workload and Constraints

ItemReported figure
Stored workflow definitions10 million
Workflow runsAbout 1 million per day
Tasks per runAverage 10: 8 I/O tasks and 2 LLM tasks
Separately stated daily task volumeAbout 100 million task executions
Peak dispatch5,000–10,000 tasks per second
I/O task duration10–50 milliseconds
LLM task duration10–60 seconds
Files and intermediate dataAbout 200 TB per day
State changes and logsAbout 100 million records, or 200 GB, per day

The run count multiplied by the average task count implies 10 million, not 100 million, daily task executions. Preserve this discrepancy as a sizing question. A state-transition record also need not correspond one-to-one with a task execution. These figures are interview inputs, not verified measurements of a deployed system.

Clarifying Questions to Ask Guidance

  • Which daily execution count is authoritative, and how many state events does one attempt generate?
  • Do task classes have the same proportions at peak as across the day?
  • What retention and query patterns apply to definitions, live run state, task history, logs, and large intermediate objects?
  • Which fairness or latency guarantees apply per tenant, and can unused capacity be borrowed?

Part 1 — Size and Separate Fast and Slow Work

Estimate average dispatch rates under both stated daily volumes and explain how runtime differences affect in-flight concurrency. Design queues and worker pools that prevent slow LLM calls from exhausting the slots needed by fast I/O tasks.

What This Part Should Cover Guidance

  • Correct arithmetic and explicit uncertainty rather than silently choosing a workload figure.
  • Separate admission, concurrency, and backpressure for task classes.
  • Capacity estimates that distinguish averages, bursts, and provider-limited throughput.

Part 2 — Store Definitions, State, Logs, and Payloads

Describe records, indexes, partition keys, retention, and storage tiers for the reported scale. Explain how to query a run's current progress without scanning its entire historical log or carrying large document bodies in the metadata database.

What This Part Should Cover Guidance

  • Immutable definition versions and durable run/task/attempt identities.
  • Hot state versus append-only history and large object payloads.
  • Partitioning, bounded indexes, retention, and recovery-safe archival.

Part 3 — Protect Tenants from Bursty Neighbors

A tenant submits hundreds of thousands of runs at once. Explain admission and scheduling controls that protect other tenants while allowing useful work to continue.

What This Part Should Cover Guidance

  • Tenant limits on submissions, queued work, running tasks, and scarce downstream resources.
  • Fair dispatch within each task class and controlled borrowing of spare capacity.
  • Metrics that expose class-specific and tenant-specific starvation.

What a Strong Answer Covers Guidance

  • Capacity reasoning that keeps the conflicting daily figures visible.
  • Isolation of fast I/O work, slow model calls, tenant demand, and downstream bottlenecks.
  • Storage and retention choices that preserve current-state reads and recovery without putting daily payload traffic in the scheduling database.

Follow-up Questions Guidance

  • How would a hot tenant change the effectiveness of sharding only by tenant ID?
  • What happens to admission when object storage or the LLM provider is slower than the workers?
  • Which state must remain available before historical events can be archived or removed?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...