Build a Scalable Text-to-Speech Service Prototype, Then Defend and Extend It
Company: Luma
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
This prompt is one of five offered in a take-home for a software engineering role on an AI product team. The candidate picks one prompt; across the set they cover frontend, backend and infrastructure work. The take-home is budgeted at 8–12 hours. The chat logs from the AI coding tools you use (such as Cursor, Codex or Claude Code) are submitted along with the code, and the work is judged on agentic development practice, 0-to-1 product scoping, and user or developer experience design. A later onsite round reviews the submission in detail.
**Prompt:** build a scalable text-to-speech (TTS) system.
Assume you may call an existing TTS model or a hosted TTS API for the synthesis itself; the work is the system around it. Confirm this before you start.
### Constraints and Clarifications
- Time box: 8–12 hours for design, code, tests and a written explanation.
- AI coding tools are allowed, and their session logs are part of what is evaluated.
- The review round treats the submission as a prototype, so anything you defer should be written down with a reason.
### Clarifying Questions
- Who is the user: developers calling an API, people using an app, or both? Is playback while audio is still being generated required, or is a finished file enough?
- Is a TTS model or provider supplied, or do you choose one? Is GPU hardware available, or only a hosted API?
- How long are inputs: single sentences, or long documents such as articles and books?
- What does "scalable" mean here: many concurrent short requests, a large volume of characters per day, or a latency target for the first audio?
- Which voices, languages and audio formats must be supported?
### Part 1 — Scope and build the prototype
Decide what the prototype must do within the time box and what it defers. Describe its API, request flow, data model, and user or developer experience, and explain how you will use AI coding agents so that the submitted logs show a disciplined workflow.
```hint Pick the scaling property to prove
Decide early which property the prototype must demonstrate: long inputs, many concurrent requests, or fast first audio. Let that choice decide what you stub out.
```
#### What This Part Should Cover
- An explicit scope cut for the time box: what ships, what is stubbed, and why.
- A request path that keeps the API responsive for both short and long inputs, including how long text is split and reassembled.
- A data model and API that support progress reporting, retries and repeated requests.
- Agentic practice visible in the logs: a written spec as context, small tasks, tests, and reviewed diffs.
### Part 2 — Product review: trade-offs and scaling limits
In the onsite review, the interviewer treats your submission as a prototype and probes it in detail; expect every shortcut to be found and asked about. Explain the trade-offs you made, where the system stops scaling first and why, which needs of TTS customers it does not yet meet, and where its UI or developer experience falls short.
```hint Measure before you shard
Estimate how much audio one synthesis worker produces per second of compute and what that costs, before you discuss caching or sharding.
```
#### What This Part Should Cover
- A back-of-the-envelope throughput and cost estimate that identifies the first bottleneck.
- The trade-offs behind chunk size, streaming versus finished files, and self-hosted versus hosted synthesis.
- Customer needs the prototype misses, such as consistent voice across chunks, pronunciation control and usage limits.
- Failure handling and UX gaps: crashed workers, repeated requests, progress and error states.
### Part 3 — A new customer requirement
The interviewer then introduces requirements from other customers and asks how you would integrate them, how long the work would take, and how it changes the roadmap. The requirements vary; practice with this one: a customer wants to submit many long documents at once through the API, be notified when each narration is finished, and hear one consistent voice across each document.
```hint Separate what unblocks from what can wait
Identify the smallest change that lets this customer start using the service, and which parts can follow later without breaking them.
```
#### What This Part Should Cover
- API and system changes for bulk submission and completion notifications, including retries and duplicate protection.
- How voice consistency is preserved across chunks and across model upgrades.
- A task-level estimate with risks, and a roadmap order with success measures.
### What a Strong Answer Covers
- A scope that fits the time box and still delivers one working end-to-end path, with deferred items stated.
- Correct handling of long inputs, repeated requests and worker failures, not only the happy path.
- Continuity from Part 1 decisions to the Part 2 limits and the Part 3 plan.
- Evidence of disciplined use of AI coding agents rather than wholesale acceptance of generated code.
### Follow-up Questions
- How would you cut the time to first audio for a long document?
- What must a cache key for synthesized audio include so that cached audio is never wrong after a change?
- A model upgrade changes how a voice sounds. How do you roll it out without breaking customers midway through a project?
- How would you meter usage and enforce per-customer quotas?
Overview: An 8 to 12 hour take-home asks you to build a scalable text-to-speech service with AI coding agents, then defend its trade-offs and scaling limits in a product review and plan a new customer requirement. Tests 0-to-1 scoping, streaming and batch synthesis design, capacity reasoning, and estimation and roadmap skills.
Read the full Luma Software Engineer interview experience this question came from