Design a Batched LLM Inference Service Assuming Each Batch Takes 100 ms

Quick Overview

A system design question on serving large language model inference by batching requests, given the interviewer's assumption that every batch finishes in 100 ms. It tests batching policy, capacity math, latency budgets, returning results to callers, failure handling, and whether a fixed batch time holds when prompt and response lengths vary.

Design a Batched LLM Inference Service Assuming Each Batch Takes 100 ms

Company: Anthropic

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Onsite

Design a service that serves inference requests for a large language model. Each client sends one request containing a prompt and waits for the model's response to that request. To use the model hardware efficiently, the service does not run requests one at a time: it groups pending requests into batches and runs each batch through the model. The interviewer fixed one simplifying assumption: the model completes each batch in 100 ms. Design the path from the client to the model and back, the policy that forms batches, and how the service scales and handles failures under that assumption. Be explicit about what the 100 ms covers and what it leaves out. ```hint Capacity from a fixed batch time If every batch takes the same 100 ms, work out how many requests one model worker can finish per second as a function of batch size, and what that tells you about how many workers you need. ``` ```hint When to close a batch Decide what sends a batch to the model: the batch filling up, a timer running out, or a worker becoming free. Each trigger trades latency against utilization differently, and the best one may depend on load. ``` ```hint One batch, many callers A batch comes back with many results at once. Decide how each result reaches the connection of the client that is still waiting for it, and what happens when that client has already given up. ``` ### Constraints and Clarifications - Treat the 100 ms batch time as given for the core design. - Request volume, the maximum batch size, the number of models and the latency target are not given. Ask for them or state your own assumptions. ### Clarifying Questions - Does the 100 ms measure the model's compute time for one batch, the time to the first generated token, or the end-to-end latency a client sees? - Is there a maximum batch size, and does the 100 ms hold for every batch size up to that limit? - Does the 100 ms hold however long the prompts and the responses in the batch are? - Is there one model, or several models whose requests cannot share a batch? - What request rate and what latency target must the service meet, and does the target apply to the first token or to the full response? - Should responses be streamed token by token, or returned whole? ### What a Strong Answer Covers - A latency budget that separates network time, queueing, the wait for a batch to start, and the 100 ms of model time - A batching policy (what closes a batch, how large it can be) with its latency and utilization trade-off - Capacity arithmetic from batch time and batch size to the number of model workers, with headroom for bursts - Routing each result in a batch back to its waiting caller, plus timeouts, cancellation and backpressure when the queue grows - Failure handling for a worker that dies mid-batch and for a request that makes every batch it joins fail - An explicit challenge to the fixed 100 ms assumption: how prompt length, response length and model choice change batch time ### Follow-up Questions - In practice a batch of long prompts with long responses takes far longer than a batch of short ones. How does the scheduler change when batch time depends on the requests inside it? - One request with a very long response holds the whole batch until it finishes. How would you stop short requests from waiting for it? - Requests target several different models. How do the queues, the batches and the worker pools change? - If the 100 ms is meant as end-to-end latency rather than model time, what is left for the model after network and queueing, and is the target achievable?

Overview: A system design question on serving large language model inference by batching requests, given the interviewer's assumption that every batch finishes in 100 ms. It tests batching policy, capacity math, latency budgets, returning results to callers, failure handling, and whether a fixed batch time holds when prompt and response lengths vary.

|Home/System Design/Anthropic
Anthropic logo
Anthropic
Oct 1, 2026
hardSoftware EngineerOnsiteSystem Design
2
0

Design a service that serves inference requests for a large language model. Each client sends one request containing a prompt and waits for the model's response to that request. To use the model hardware efficiently, the service does not run requests one at a time: it groups pending requests into batches and runs each batch through the model.

The interviewer fixed one simplifying assumption: the model completes each batch in 100 ms. Design the path from the client to the model and back, the policy that forms batches, and how the service scales and handles failures under that assumption. Be explicit about what the 100 ms covers and what it leaves out.

Constraints and Clarifications

  • Treat the 100 ms batch time as given for the core design.
  • Request volume, the maximum batch size, the number of models and the latency target are not given. Ask for them or state your own assumptions.

Clarifying Questions Guidance

  • Does the 100 ms measure the model's compute time for one batch, the time to the first generated token, or the end-to-end latency a client sees?
  • Is there a maximum batch size, and does the 100 ms hold for every batch size up to that limit?
  • Does the 100 ms hold however long the prompts and the responses in the batch are?
  • Is there one model, or several models whose requests cannot share a batch?
  • What request rate and what latency target must the service meet, and does the target apply to the first token or to the full response?
  • Should responses be streamed token by token, or returned whole?

What a Strong Answer Covers Guidance

  • A latency budget that separates network time, queueing, the wait for a batch to start, and the 100 ms of model time
  • A batching policy (what closes a batch, how large it can be) with its latency and utilization trade-off
  • Capacity arithmetic from batch time and batch size to the number of model workers, with headroom for bursts
  • Routing each result in a batch back to its waiting caller, plus timeouts, cancellation and backpressure when the queue grows
  • Failure handling for a worker that dies mid-batch and for a request that makes every batch it joins fail
  • An explicit challenge to the fixed 100 ms assumption: how prompt length, response length and model choice change batch time

Follow-up Questions Guidance

  • In practice a batch of long prompts with long responses takes far longer than a batch of short ones. How does the scheduler change when batch time depends on the requests inside it?
  • One request with a very long response holds the whole batch until it finishes. How would you stop short requests from waiting for it?
  • Requests target several different models. How do the queues, the batches and the worker pools change?
  • If the 100 ms is meant as end-to-end latency rather than model time, what is left for the model after network and queueing, and is the target achievable?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...