Design ChatGPT: Streaming Replies over SSE and Rendering Long Conversations
Company: OpenAI
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Onsite
Design ChatGPT: a web chat application in which a user sends a message and a large language model's reply appears token by token as it is generated, within a multi-turn conversation.
The interviewer set two simplifications at the start: do not consider scale, and nothing needs to be persisted. Later, the interviewer asked why the design uses Server-Sent Events (SSE) rather than WebSockets, and how the frontend can be optimized so that long conversations still render well.
### Constraints and Clarifications
- No scale requirements: design for a single deployment and focus on the request flow, the streaming path and the client.
- No persistence: there is no database of users or conversation history. Decide where the conversation lives while it is in use.
- Unless the interviewer asks for it, treat the model as an existing inference service that accepts a prompt and streams generated tokens back; you design the chat system around it, not the model serving itself.
### Clarifying Questions
- Without persistence, must a conversation survive a page refresh, or may it be lost?
- Which features are in scope beyond sending a message and streaming the reply: stopping a reply, regenerating it, editing an earlier message?
- What should happen when a conversation grows longer than the model's context window?
- Do replies contain rich content, such as Markdown and code blocks, that the client must render?
### Part 1 — Core design
Design the components and the API, and walk through what happens from the moment the user sends a message until the full reply is displayed. Explain where the conversation history is kept and how it reaches the model on each turn.
```hint Where the transcript lives
With nothing persisted, either the client or the server must hold the conversation between turns. Compare what each choice means for which server can handle the next message.
```
```hint The user changes their mind
Consider what should happen on the server when the user stops a reply halfway, or closes the tab, while the model is still generating.
```
#### What This Part Should Cover
- Components and a request and response contract for sending a message and receiving a streamed reply
- Where conversation state lives under the no-persistence constraint, and how the prompt is assembled within the context window
- Cancellation, and failures before and during the stream
- The latency the user perceives, especially the time to the first token
### Part 2 — SSE or WebSockets
Your design streams the reply with Server-Sent Events. Why SSE rather than WebSockets?
```hint Match the transport to the traffic
Look at which direction data flows during one turn of the conversation, and how often the client needs to send anything while a reply is streaming.
```
```hint Everything between client and server
Consider the load balancers, proxies and HTTP tooling the stream passes through, and what each transport asks of them.
```
#### What This Part Should Cover
- The communication pattern of a chat turn and which transport it fits
- Operational differences: HTTP infrastructure, authentication, load balancing and reconnection
- Practical limitations of SSE in browsers and through proxies, and how to work around them
- When WebSockets would be the better choice
### Part 3 — Rendering long conversations
A conversation can grow very long. How would you optimize the frontend so that long conversations keep rendering smoothly, including while a new reply is streaming in?
```hint Count the work per token
Estimate how much of the page is re-rendered and re-parsed each time one token arrives, and how that grows with the length of the conversation.
```
```hint Not everything is on screen
Only a small part of a long conversation is visible at any moment. Think about what the browser is still doing for the rest.
```
#### What This Part Should Cover
- Where the rendering cost comes from as the conversation and the streaming reply grow
- Limiting re-rendering and re-parsing to what actually changed
- Reducing the cost of messages outside the viewport, and the side effects of doing so
- Scroll behavior during streaming, and how performance is measured
### What a Strong Answer Covers
- A design that respects the stated constraints instead of spending time on scale and storage
- An end-to-end streaming path with clear event semantics, cancellation and error handling
- A transport choice justified by the traffic pattern and the infrastructure, with its limitations acknowledged
- Frontend techniques tied to where the cost actually comes from, with their trade-offs
- Handling of conversations that exceed the model's context window
### Follow-up Questions
- If persistence becomes a requirement later, what would you store, and when would a streamed reply be written?
- The network drops halfway through a reply. What does the user see, and how could the reply be resumed rather than regenerated?
- For which product features would you switch to WebSockets, and what would that change in the backend?
Overview: Design a ChatGPT-style chat application without scale or persistence requirements, covering how a user message reaches the model and how the reply streams back token by token. Probes why Server-Sent Events suit this streaming better than WebSockets and how the frontend keeps long conversations rendering smoothly.
Read the full OpenAI Software Engineer interview experience this question came from