Transformer Attention: Computing Q, K and V, and How a KV Cache Speeds Up Decoding
Company: Netflix
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
In a machine learning fundamentals screen, the interviewer asked two related questions about transformer models: what attention is, including exactly how the queries, keys and values are computed, and what a KV cache is.
### Clarifying Questions
- Should the answer focus on self-attention in a decoder-only language model, or also cover cross-attention in an encoder-decoder model?
- Is a single attention head enough to start with, or should the answer go straight to multi-head attention?
- How much depth is expected: formulas with matrix shapes, or the intuition plus the key equation?
### Part 1 — Attention and the Q, K, V computation
What is attention? Starting from a sequence of $n$ token representations of width $d_{\text{model}}$, show how the queries, keys and values are computed, how they combine into the attention output, and how this extends to multi-head attention. State the shape of every matrix and explain any scaling or masking you apply.
```hint Follow the shapes
Write the input as an $n \times d_{\text{model}}$ matrix and track the shape of every product until you are back to one vector per token.
```
```hint Size of a dot product
Think about how the typical magnitude of a dot product changes as the vector dimension grows, and what very large inputs do to a softmax.
```
#### What This Part Should Cover
- The learned projections that produce queries, keys and values, with their shapes
- Scores, scaling, the softmax over keys and the weighted sum of values
- Causal masking in a decoder, and multi-head attention with its output projection
- The cost of attention as a function of sequence length
### Part 2 — The KV cache
What is a KV cache? Explain what is cached, why it is valid to reuse it, how it changes the work done for each newly generated token, and what it costs in memory.
```hint What changes between steps
When the model generates the next token, compare the keys and values it needs with the ones it already computed at the previous step. Ask which of them could possibly have changed.
```
```hint Put a number on it
Estimate the cache size for a concrete model shape and context length before deciding whether it is cheap.
```
#### What This Part Should Cover
- Why past keys and values can be reused under causal masking, and why past queries are not needed
- The prefill and decode phases, and the per-token compute saved
- Cache memory as a function of layers, heads, head dimension, sequence length, batch size and precision
- Techniques that shrink or manage the cache, and their trade-offs
### What a Strong Answer Covers
- Correct attention formulas with consistent matrix shapes
- An explicit link between causal masking and the validity of the cache
- A quantified trade-off: compute saved per token versus memory spent
- The serving consequences of the cache, such as limits on batch size and context length
### Follow-up Questions
- Why can an encoder-only model with bidirectional attention not reuse a KV cache when a token is appended, the way a decoder does?
- How do multi-query and grouped-query attention reduce the cache size, and what do they give up?
- Why is token-by-token decoding usually limited by memory bandwidth rather than by arithmetic, even with a KV cache?
- With very long contexts, the cache no longer fits in GPU memory. What options do you have?
Overview: Explain how transformer attention works: computing queries, keys and values from token representations, scaling the scores, applying a causal mask and combining heads. Then describe the KV cache used in autoregressive decoding, why reusing past keys and values is valid, and what it saves in compute and costs in memory.
Read the full Netflix Machine Learning Engineer interview experience this question came from