L1 vs L2 Normalization: Definitions, Geometry and When to Use Each

Read the full interview experience this question came from →

Quick Overview

A machine learning fundamentals question asking how L1 normalization differs from L2 normalization of a vector. It tests precise definitions, what each normalized vector represents geometrically, when each is the right choice, and the ability to keep normalization distinct from L1 and L2 regularization.

L1 vs L2 Normalization: Definitions, Geometry and When to Use Each

Company: Netflix

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: hard

Interview Round: Technical Screen

What is the difference between L1 normalization and L2 normalization? This was asked as a short machine learning fundamentals question. Answer it as you would aloud: define both, show what each does to a vector, and explain when you would choose one over the other. ```hint Divide by what Write down the norm each method divides by, and the quantity that is guaranteed to equal one afterwards. ``` ```hint What the output is used for Think about what the dot product of two normalized vectors means in each case, and what a non-negative vector becomes after each kind of normalization. ``` ### Clarifying Questions - Do you mean rescaling a vector by its L1 or L2 norm, or the L1 and L2 regularization penalties on model weights? The terms are often mixed up, so confirm which one is meant, or cover both and label them clearly. - Is the normalization applied to each sample (a row of the feature matrix) or to each feature (a column)? ### What a Strong Answer Covers - Precise definitions of both norms and of each normalization - A concrete numeric example that shows how the results differ - The geometric meaning of the normalized vectors and the downstream use each one supports - How each behaves with one dominant component, with negative values and with the zero vector - A clear separation from L1 and L2 regularization and from per-feature standardization ### Follow-up Questions - You L2-normalize embeddings before a nearest-neighbor search. Why does ranking by dot product then give the same order as ranking by cosine similarity and by Euclidean distance? - How do L1 and L2 regularization differ in their effect on model weights, and why does only L1 tend to produce exact zeros? - How does L2 normalization relate to the RMSNorm and LayerNorm layers used inside neural networks? - When would per-sample normalization hurt a model rather than help it?

Overview: A machine learning fundamentals question asking how L1 normalization differs from L2 normalization of a vector. It tests precise definitions, what each normalized vector represents geometrically, when each is the right choice, and the ability to keep normalization distinct from L1 and L2 regularization.

Read the full Netflix Machine Learning Engineer interview experience this question came from

|Home/Machine Learning/Netflix
Netflix logo
Netflix
Sep 28, 2026
hardMachine Learning EngineerTechnical ScreenMachine Learning
0
0

What is the difference between L1 normalization and L2 normalization?

This was asked as a short machine learning fundamentals question. Answer it as you would aloud: define both, show what each does to a vector, and explain when you would choose one over the other.

Clarifying Questions Guidance

  • Do you mean rescaling a vector by its L1 or L2 norm, or the L1 and L2 regularization penalties on model weights? The terms are often mixed up, so confirm which one is meant, or cover both and label them clearly.
  • Is the normalization applied to each sample (a row of the feature matrix) or to each feature (a column)?

What a Strong Answer Covers Guidance

  • Precise definitions of both norms and of each normalization
  • A concrete numeric example that shows how the results differ
  • The geometric meaning of the normalized vectors and the downstream use each one supports
  • How each behaves with one dominant component, with negative values and with the zero vector
  • A clear separation from L1 and L2 regularization and from per-feature standardization

Follow-up Questions Guidance

  • You L2-normalize embeddings before a nearest-neighbor search. Why does ranking by dot product then give the same order as ranking by cosine similarity and by Euclidean distance?
  • How do L1 and L2 regularization differ in their effect on model weights, and why does only L1 tend to produce exact zeros?
  • How does L2 normalization relate to the RMSNorm and LayerNorm layers used inside neural networks?
  • When would per-sample normalization hurt a model rather than help it?
Loading comments...