Implement multi-head self-attention from scratch with projections and masking

Quick Overview

Implement a multi-head self-attention layer from scratch, with query, key, value and output projections and an optional attention mask, returning a tensor shaped like its input. Tests tensor shape handling, masking semantics, numerical stability and the cost of attention as sequence length grows.

Implement multi-head self-attention from scratch with projections and masking

Company: ByteDance

Role: Software Engineer

Category: Machine Learning

Difficulty: hard

Interview Round: Technical Screen

Implement a multi-head self-attention layer, the attention block used in Transformer models, from scratch. The interview allowed about 25 minutes for it. The layer receives an input batch `x` of shape `(batch, seq_len, d_model)` and a number of heads `num_heads`, and returns a tensor of the same shape. Include the learned query, key, value and output projections, and support an optional attention mask. ```hint Track every shape Write down the tensor shape after each step, especially where the model dimension is split into heads and where the heads are merged back. ``` ```hint Where the mask belongs Decide at which step the mask must be applied so that a masked key receives exactly zero weight, and what should happen when every key in a row is masked. ``` ### Constraints and Clarifications - Use PyTorch tensor operations or NumPy. Do not call a built-in attention module or a fused attention function. - `d_model` is divisible by `num_heads`; each head works in dimension `d_k = d_model // num_heads`. - Assume the mask is boolean and broadcastable to `(batch, num_heads, seq_len, seq_len)`, with `True` marking query-key pairs that may attend. Libraries use opposite conventions, so confirm this with the interviewer. - Dropout on the attention weights is optional. ### Clarifying Questions - Is this self-attention only, or should the layer accept separate query and key/value inputs (cross-attention)? - Which masks are needed: a padding mask, a causal mask, or both? - Should the layer also return the attention weights? - Is a trainable module expected, or is a forward pass with given weight matrices enough? ### What a Strong Answer Covers - Correct projections and a split into heads that keeps positions and features of different heads apart - Scaled dot-product scores with the right scaling factor and a numerically stable softmax over the correct axis - Correct masking semantics, including rows in which every key is masked - Merging the heads and applying the output projection, with the expected output shape - Time and memory cost as a function of sequence length, and a quick way to test the layer ### Follow-up Questions - Why are the scores divided by $\sqrt{d_k}$, and what happens to training without it? - How would you add a causal mask, and how would you cache keys and values for token-by-token decoding? - What limits this layer on very long sequences, and what would you change? - How does grouped-query or multi-query attention change your implementation?

Overview: Implement a multi-head self-attention layer from scratch, with query, key, value and output projections and an optional attention mask, returning a tensor shaped like its input. Tests tensor shape handling, masking semantics, numerical stability and the cost of attention as sequence length grows.

|Home/Machine Learning/ByteDance
ByteDance logo
ByteDance
Oct 4, 2026
hardSoftware EngineerTechnical ScreenMachine Learning
0
0

Implement a multi-head self-attention layer, the attention block used in Transformer models, from scratch. The interview allowed about 25 minutes for it.

The layer receives an input batch x of shape (batch, seq_len, d_model) and a number of heads num_heads, and returns a tensor of the same shape. Include the learned query, key, value and output projections, and support an optional attention mask.

Constraints and Clarifications

  • Use PyTorch tensor operations or NumPy. Do not call a built-in attention module or a fused attention function.
  • d_model is divisible by num_heads ; each head works in dimension d_k = d_model // num_heads .
  • Assume the mask is boolean and broadcastable to (batch, num_heads, seq_len, seq_len) , with True marking query-key pairs that may attend. Libraries use opposite conventions, so confirm this with the interviewer.
  • Dropout on the attention weights is optional.

Clarifying Questions Guidance

  • Is this self-attention only, or should the layer accept separate query and key/value inputs (cross-attention)?
  • Which masks are needed: a padding mask, a causal mask, or both?
  • Should the layer also return the attention weights?
  • Is a trainable module expected, or is a forward pass with given weight matrices enough?

What a Strong Answer Covers Guidance

  • Correct projections and a split into heads that keeps positions and features of different heads apart
  • Scaled dot-product scores with the right scaling factor and a numerically stable softmax over the correct axis
  • Correct masking semantics, including rows in which every key is masked
  • Merging the heads and applying the output projection, with the expected output shape
  • Time and memory cost as a function of sequence length, and a quick way to test the layer

Follow-up Questions Guidance

  • Why are the scores divided by dk\sqrt{d_k} , and what happens to training without it?
  • How would you add a causal mask, and how would you cache keys and values for token-by-token decoding?
  • What limits this layer on very long sequences, and what would you change?
  • How does grouped-query or multi-query attention change your implementation?
Loading comments...