Implement multi-head self-attention from scratch with projections and masking
Company: ByteDance
Role: Software Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
Implement a multi-head self-attention layer, the attention block used in Transformer models, from scratch. The interview allowed about 25 minutes for it.
The layer receives an input batch `x` of shape `(batch, seq_len, d_model)` and a number of heads `num_heads`, and returns a tensor of the same shape. Include the learned query, key, value and output projections, and support an optional attention mask.
```hint Track every shape
Write down the tensor shape after each step, especially where the model dimension is split into heads and where the heads are merged back.
```
```hint Where the mask belongs
Decide at which step the mask must be applied so that a masked key receives exactly zero weight, and what should happen when every key in a row is masked.
```
### Constraints and Clarifications
- Use PyTorch tensor operations or NumPy. Do not call a built-in attention module or a fused attention function.
- `d_model` is divisible by `num_heads`; each head works in dimension `d_k = d_model // num_heads`.
- Assume the mask is boolean and broadcastable to `(batch, num_heads, seq_len, seq_len)`, with `True` marking query-key pairs that may attend. Libraries use opposite conventions, so confirm this with the interviewer.
- Dropout on the attention weights is optional.
### Clarifying Questions
- Is this self-attention only, or should the layer accept separate query and key/value inputs (cross-attention)?
- Which masks are needed: a padding mask, a causal mask, or both?
- Should the layer also return the attention weights?
- Is a trainable module expected, or is a forward pass with given weight matrices enough?
### What a Strong Answer Covers
- Correct projections and a split into heads that keeps positions and features of different heads apart
- Scaled dot-product scores with the right scaling factor and a numerically stable softmax over the correct axis
- Correct masking semantics, including rows in which every key is masked
- Merging the heads and applying the output projection, with the expected output shape
- Time and memory cost as a function of sequence length, and a quick way to test the layer
### Follow-up Questions
- Why are the scores divided by $\sqrt{d_k}$, and what happens to training without it?
- How would you add a causal mask, and how would you cache keys and values for token-by-token decoding?
- What limits this layer on very long sequences, and what would you change?
- How does grouped-query or multi-query attention change your implementation?
Overview: Implement a multi-head self-attention layer from scratch, with query, key, value and output projections and an optional attention mask, returning a tensor shaped like its input. Tests tensor shape handling, masking semantics, numerical stability and the cost of attention as sequence length grows.