Transformer Attention: Computing Q, K and V, and How a KV Cache Speeds Up Decoding

Read the full interview experience this question came from →

Quick Overview

Explain how transformer attention works: computing queries, keys and values from token representations, scaling the scores, applying a causal mask and combining heads. Then describe the KV cache used in autoregressive decoding, why reusing past keys and values is valid, and what it saves in compute and costs in memory.

Transformer Attention: Computing Q, K and V, and How a KV Cache Speeds Up Decoding

Company: Netflix

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: hard

Interview Round: Technical Screen

In a machine learning fundamentals screen, the interviewer asked two related questions about transformer models: what attention is, including exactly how the queries, keys and values are computed, and what a KV cache is. ### Clarifying Questions - Should the answer focus on self-attention in a decoder-only language model, or also cover cross-attention in an encoder-decoder model? - Is a single attention head enough to start with, or should the answer go straight to multi-head attention? - How much depth is expected: formulas with matrix shapes, or the intuition plus the key equation? ### Part 1 — Attention and the Q, K, V computation What is attention? Starting from a sequence of $n$ token representations of width $d_{\text{model}}$, show how the queries, keys and values are computed, how they combine into the attention output, and how this extends to multi-head attention. State the shape of every matrix and explain any scaling or masking you apply. ```hint Follow the shapes Write the input as an $n \times d_{\text{model}}$ matrix and track the shape of every product until you are back to one vector per token. ``` ```hint Size of a dot product Think about how the typical magnitude of a dot product changes as the vector dimension grows, and what very large inputs do to a softmax. ``` #### What This Part Should Cover - The learned projections that produce queries, keys and values, with their shapes - Scores, scaling, the softmax over keys and the weighted sum of values - Causal masking in a decoder, and multi-head attention with its output projection - The cost of attention as a function of sequence length ### Part 2 — The KV cache What is a KV cache? Explain what is cached, why it is valid to reuse it, how it changes the work done for each newly generated token, and what it costs in memory. ```hint What changes between steps When the model generates the next token, compare the keys and values it needs with the ones it already computed at the previous step. Ask which of them could possibly have changed. ``` ```hint Put a number on it Estimate the cache size for a concrete model shape and context length before deciding whether it is cheap. ``` #### What This Part Should Cover - Why past keys and values can be reused under causal masking, and why past queries are not needed - The prefill and decode phases, and the per-token compute saved - Cache memory as a function of layers, heads, head dimension, sequence length, batch size and precision - Techniques that shrink or manage the cache, and their trade-offs ### What a Strong Answer Covers - Correct attention formulas with consistent matrix shapes - An explicit link between causal masking and the validity of the cache - A quantified trade-off: compute saved per token versus memory spent - The serving consequences of the cache, such as limits on batch size and context length ### Follow-up Questions - Why can an encoder-only model with bidirectional attention not reuse a KV cache when a token is appended, the way a decoder does? - How do multi-query and grouped-query attention reduce the cache size, and what do they give up? - Why is token-by-token decoding usually limited by memory bandwidth rather than by arithmetic, even with a KV cache? - With very long contexts, the cache no longer fits in GPU memory. What options do you have?

Overview: Explain how transformer attention works: computing queries, keys and values from token representations, scaling the scores, applying a causal mask and combining heads. Then describe the KV cache used in autoregressive decoding, why reusing past keys and values is valid, and what it saves in compute and costs in memory.

Read the full Netflix Machine Learning Engineer interview experience this question came from

|Home/Machine Learning/Netflix
Netflix logo
Netflix
Sep 28, 2026
hardMachine Learning EngineerTechnical ScreenMachine Learning
0
0

In a machine learning fundamentals screen, the interviewer asked two related questions about transformer models: what attention is, including exactly how the queries, keys and values are computed, and what a KV cache is.

Clarifying Questions Guidance

  • Should the answer focus on self-attention in a decoder-only language model, or also cover cross-attention in an encoder-decoder model?
  • Is a single attention head enough to start with, or should the answer go straight to multi-head attention?
  • How much depth is expected: formulas with matrix shapes, or the intuition plus the key equation?

Part 1 — Attention and the Q, K, V computation

What is attention? Starting from a sequence of nn token representations of width dmodeld_{\text{model}}, show how the queries, keys and values are computed, how they combine into the attention output, and how this extends to multi-head attention. State the shape of every matrix and explain any scaling or masking you apply.

What This Part Should Cover Guidance

  • The learned projections that produce queries, keys and values, with their shapes
  • Scores, scaling, the softmax over keys and the weighted sum of values
  • Causal masking in a decoder, and multi-head attention with its output projection
  • The cost of attention as a function of sequence length

Part 2 — The KV cache

What is a KV cache? Explain what is cached, why it is valid to reuse it, how it changes the work done for each newly generated token, and what it costs in memory.

What This Part Should Cover Guidance

  • Why past keys and values can be reused under causal masking, and why past queries are not needed
  • The prefill and decode phases, and the per-token compute saved
  • Cache memory as a function of layers, heads, head dimension, sequence length, batch size and precision
  • Techniques that shrink or manage the cache, and their trade-offs

What a Strong Answer Covers Guidance

  • Correct attention formulas with consistent matrix shapes
  • An explicit link between causal masking and the validity of the cache
  • A quantified trade-off: compute saved per token versus memory spent
  • The serving consequences of the cache, such as limits on batch size and context length

Follow-up Questions Guidance

  • Why can an encoder-only model with bidirectional attention not reuse a KV cache when a token is appended, the way a decoder does?
  • How do multi-query and grouped-query attention reduce the cache size, and what do they give up?
  • Why is token-by-token decoding usually limited by memory bandwidth rather than by arithmetic, even with a KV cache?
  • With very long contexts, the cache no longer fits in GPU memory. What options do you have?
Loading comments...