L1 vs L2 Normalization: Definitions, Geometry and When to Use Each
Company: Netflix
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
What is the difference between L1 normalization and L2 normalization?
This was asked as a short machine learning fundamentals question. Answer it as you would aloud: define both, show what each does to a vector, and explain when you would choose one over the other.
```hint Divide by what
Write down the norm each method divides by, and the quantity that is guaranteed to equal one afterwards.
```
```hint What the output is used for
Think about what the dot product of two normalized vectors means in each case, and what a non-negative vector becomes after each kind of normalization.
```
### Clarifying Questions
- Do you mean rescaling a vector by its L1 or L2 norm, or the L1 and L2 regularization penalties on model weights? The terms are often mixed up, so confirm which one is meant, or cover both and label them clearly.
- Is the normalization applied to each sample (a row of the feature matrix) or to each feature (a column)?
### What a Strong Answer Covers
- Precise definitions of both norms and of each normalization
- A concrete numeric example that shows how the results differ
- The geometric meaning of the normalized vectors and the downstream use each one supports
- How each behaves with one dominant component, with negative values and with the zero vector
- A clear separation from L1 and L2 regularization and from per-feature standardization
### Follow-up Questions
- You L2-normalize embeddings before a nearest-neighbor search. Why does ranking by dot product then give the same order as ranking by cosine similarity and by Euclidean distance?
- How do L1 and L2 regularization differ in their effect on model weights, and why does only L1 tend to produce exact zeros?
- How does L2 normalization relate to the RMSNorm and LayerNorm layers used inside neural networks?
- When would per-sample normalization hurt a model rather than help it?
Overview: A machine learning fundamentals question asking how L1 normalization differs from L2 normalization of a vector. It tests precise definitions, what each normalized vector represents geometrically, when each is the right choice, and the ability to keep normalization distinct from L1 and L2 regularization.
Read the full Netflix Machine Learning Engineer interview experience this question came from