Design Streaming Speech and Multi-Model Chat Systems

Read the full interview experience this question came from →

Quick Overview

A two-part system-design interview covering incremental multi-speaker transcription and a streaming multi-model chat service. Candidates must reason about ordering, revisions, backpressure, model fallback, durable conversation state, safety, and recovery.

Design Streaming Speech and Multi-Model Chat Systems

Company: Distyl

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

# Design Streaming Speech and Multi-Model Chat Systems Redesign a batch speech-transcription pipeline to process audio incrementally and support multiple speakers. Then design an LLM chat service that streams responses, switches models when rate or token limits are reached, scales behind a load balancer, and stores conversation history safely. ### Constraints & Assumptions - Audio chunks can arrive late, be duplicated, or end mid-utterance. - Partial transcripts may be revised as more context arrives. - A model switch must not silently change the request's safety or tool permissions. - Load balancers cannot rely on one process retaining all stream state. ### Clarifying Questions to Ask - What latency and final-transcript accuracy are required? - Is speaker diarization based on separate channels or mixed audio? - Which provider limits are per request, model, tenant, or time window? ### Part 1 — Streaming transcription Define chunking, ordering, buffering, partial and final transcript states, speaker handling, backpressure, and performance measurement. #### What This Part Should Cover - Sequence IDs and bounded reordering - Revision-aware transcript segments - Diarization, VAD, latency, and throughput diagnostics ### Part 2 — Chat serving Design request routing, streaming through a load balancer, model fallback, rate-limit handling, and conversation persistence. #### What This Part Should Cover - Externalized stream and request state - Capability-aware routing and idempotency - Durable message lifecycle and token accounting ### What a Strong Answer Covers - Ordering and recovery semantics - Backpressure and cancellation - Consistent safety, authorization, and storage across model switches ```hint Make partial output a versioned state A partial transcript or model response is not an append-only final record. Give each segment or message a status and revision so clients can replace provisional text. ``` ### Follow-up Questions - How would you recover a stream after a worker dies? - When should conversation content be written to the database?

Overview: A two-part system-design interview covering incremental multi-speaker transcription and a streaming multi-model chat service. Candidates must reason about ordering, revisions, backpressure, model fallback, durable conversation state, safety, and recovery.

Read the full Distyl Machine Learning Engineer interview experience this question came from

|Home/ML System Design/Distyl
Distyl logo
Distyl
May 1, 2026
mediumMachine Learning EngineerOnsiteML System Design
6
0

Design Streaming Speech and Multi-Model Chat Systems

Redesign a batch speech-transcription pipeline to process audio incrementally and support multiple speakers. Then design an LLM chat service that streams responses, switches models when rate or token limits are reached, scales behind a load balancer, and stores conversation history safely.

Constraints & Assumptions

  • Audio chunks can arrive late, be duplicated, or end mid-utterance.
  • Partial transcripts may be revised as more context arrives.
  • A model switch must not silently change the request's safety or tool permissions.
  • Load balancers cannot rely on one process retaining all stream state.

Clarifying Questions to Ask Guidance

  • What latency and final-transcript accuracy are required?
  • Is speaker diarization based on separate channels or mixed audio?
  • Which provider limits are per request, model, tenant, or time window?

Part 1 — Streaming transcription

Define chunking, ordering, buffering, partial and final transcript states, speaker handling, backpressure, and performance measurement.

What This Part Should Cover Guidance

  • Sequence IDs and bounded reordering
  • Revision-aware transcript segments
  • Diarization, VAD, latency, and throughput diagnostics

Part 2 — Chat serving

Design request routing, streaming through a load balancer, model fallback, rate-limit handling, and conversation persistence.

What This Part Should Cover Guidance

  • Externalized stream and request state
  • Capability-aware routing and idempotency
  • Durable message lifecycle and token accounting

What a Strong Answer Covers Guidance

  • Ordering and recovery semantics
  • Backpressure and cancellation
  • Consistent safety, authorization, and storage across model switches

Follow-up Questions Guidance

  • How would you recover a stream after a worker dies?
  • When should conversation content be written to the database?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...