Explain LLM generational improvements, multimodal LLMs, and normalization in LLMs
Company: ByteDance
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
The interviewer asked three conceptual questions about large language models (LLMs): how models have improved from one generation to the next according to their technical reports, how multimodal models work, and why LLMs use the normalization they do. Answer each part.
### Constraints and Clarifications
- The third question was reported as asking why LLMs now commonly use "linear normalization". The standard reading is layer normalization and its variants, such as RMSNorm; confirm the intended term with the interviewer if it is unclear.
- Claims about specific models should stay within what their public reports state.
### Clarifying Questions
- In Part 1, may I choose the model family I know best, and should I focus on architecture, data, training recipe, or all of them?
- In Part 2, which modalities matter: images and text only, or also audio and video?
- In Part 3, is the question about layer normalization versus batch normalization, about RMSNorm versus standard layer normalization, or about where normalization sits in the transformer block?
### Part 1 — Improvements across model generations
Have you read the technical reports of different LLMs? Pick a model family and explain the main improvements each generation made over the previous one.
```hint Sort the changes
Group the generational changes into data, architecture, training recipe, post-training and inference efficiency, then say which ones mattered most and why each was made.
```
#### What This Part Should Cover
- Concrete changes per generation in data, architecture, post-training and efficiency
- The motivation for each change, not only its name
- Which gains came from scale and which from changes in method
- Care to separate what the reports state from speculation
### Part 2 — Multimodal models
Explain how a multimodal LLM that accepts images together with text is typically built and trained.
```hint From pixels to the input sequence
The language model consumes a sequence of vectors. Think about what turns an image into such vectors and what makes them compatible with the text embeddings the model already understands.
```
#### What This Part Should Cover
- The vision encoder, and the connector that maps visual features into the language model's input space
- Alternative fusion designs and their trade-offs
- The usual training stages and which components are frozen in each
- Practical costs and failure modes, such as visual token count, image resolution and hallucinated details
### Part 3 — Normalization in LLMs
Why do current LLMs normalize activations with layer normalization or its variants rather than with batch normalization, and why is the normalization usually placed before each sublayer of the transformer block?
```hint What the statistics are computed over
Compare the dimensions each scheme averages over, then consider variable-length sequences, small per-device batches, causal masking and generating one token at a time.
```
```hint Follow the residual path
Consider how gradients travel through a very deep stack of blocks when normalization sits on the residual path versus inside each branch.
```
#### What This Part Should Cover
- Why batch statistics fit poorly with autoregressive sequence models
- What layer normalization and RMSNorm compute, and why RMSNorm can drop mean-centering
- Pre-norm versus post-norm placement and its effect on training stability
- The compute cost of normalization during training and inference
### What a Strong Answer Covers
- Technically accurate explanations that state why each design choice was made
- Specific, verifiable examples from public model reports instead of vague trends
- Clear trade-offs for each choice: fusion design, frozen versus trained components, normalization type and placement
- Links between the parts, such as efficiency-driven architecture changes that also shape multimodal and long-context models
### Follow-up Questions
- Pre-norm transformers train stably but are sometimes said to make poor use of their deepest layers. What would you try to keep the stability without that cost?
- How does grouped-query attention reduce inference cost, and what does it give up compared with standard multi-head attention?
- How would you handle high-resolution images containing small text without an explosion in the number of visual tokens?
- How would you test whether a multimodal model actually uses the image rather than guessing from the text?
Overview: Three conceptual LLM questions: the main improvements across generations of a model family as described in technical reports, how multimodal image-and-text models are built and trained, and why modern LLMs use layer normalization or RMSNorm with pre-norm placement instead of batch normalization.
Read the full ByteDance Machine Learning Engineer interview experience this question came from