What Happens When All Neural-Network Weights Start at Zero?
Company: Snapchat
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
# What Happens When All Neural-Network Weights Start at Zero?
What happens if every weight in a neural network is initialized to zero? Explain why the answer depends on the architecture and activation functions. Focus on a multilayer network with several interchangeable hidden units, and state the bias initialization when using that example. Contrast the result with a linear or logistic model that has no hidden layer.
Explain the learning limitation caused by identically initialized hidden units and discuss whether all gradients must be zero in every possible network. Your answer should distinguish failure to break symmetry from the specific zero-gradient behavior of a particular activation and loss implementation.
### What a Strong Answer Covers
- Why identical hidden units can receive identical updates and fail to learn distinct representations.
- How zero activations or zero downstream weights can suppress gradients in a specific multilayer example.
- The fact that zero initialization is not universally invalid for models without hidden-unit symmetry.
- A practical initialization strategy and why biases need not follow the same rule as weights.
### Follow-up Questions
- Could zero hidden-layer weights with different random biases break some symmetry?
- Why is initializing all hidden weights to the same nonzero constant also problematic?
Overview: Explain zero-weight initialization through hidden-unit symmetry and gradient flow, and contrast it with linear and logistic models.
What Happens When All Neural-Network Weights Start at Zero?
What happens if every weight in a neural network is initialized to zero? Explain why the answer depends on the architecture and activation functions. Focus on a multilayer network with several interchangeable hidden units, and state the bias initialization when using that example. Contrast the result with a linear or logistic model that has no hidden layer.
Explain the learning limitation caused by identically initialized hidden units and discuss whether all gradients must be zero in every possible network. Your answer should distinguish failure to break symmetry from the specific zero-gradient behavior of a particular activation and loss implementation.
What a Strong Answer Covers Guidance
Why identical hidden units can receive identical updates and fail to learn distinct representations.
How zero activations or zero downstream weights can suppress gradients in a specific multilayer example.
The fact that zero initialization is not universally invalid for models without hidden-unit symmetry.
A practical initialization strategy and why biases need not follow the same rule as weights.
Follow-up Questions Guidance
Could zero hidden-layer weights with different random biases break some symmetry?
Why is initializing all hidden weights to the same nonzero constant also problematic?