Explain Vanishing Gradients in Neural-Network Training
Company: Snapchat
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
# Explain Vanishing Gradients in Neural-Network Training
What is the vanishing-gradient problem, why can it occur when training a neural network, and what would you do to diagnose and reduce it? Explain how gradients propagate through repeated layers or time steps and distinguish very small backpropagated gradients from a model that simply has a low loss near a useful optimum.
You may illustrate the mechanism with a deep feed-forward network or a recurrent network, but state which setting you are using. No architecture or activation function is specified, so explain which proposed remedies depend on those choices. Discuss why a change that helps gradient propagation is not by itself evidence that the trained model generalizes better.
### What a Strong Answer Covers
- The role of the chain rule and repeated Jacobian products in shrinking earlier-layer gradients.
- Saturating activations, initialization and scale as possible contributors.
- Layer-wise gradient and activation diagnostics, plus checks for disconnected computation or implementation errors.
- Architecture-appropriate mitigations and evaluation after changing the training setup.
### Follow-up Questions
- Why does clipping gradients usually address a different problem?
- Can a rectified activation still leave a unit unable to learn?
Overview: Diagnose vanishing gradients through the chain rule, activation and gradient checks, and architecture-appropriate training remedies.
Explain Vanishing Gradients in Neural-Network Training
What is the vanishing-gradient problem, why can it occur when training a neural network, and what would you do to diagnose and reduce it? Explain how gradients propagate through repeated layers or time steps and distinguish very small backpropagated gradients from a model that simply has a low loss near a useful optimum.
You may illustrate the mechanism with a deep feed-forward network or a recurrent network, but state which setting you are using. No architecture or activation function is specified, so explain which proposed remedies depend on those choices. Discuss why a change that helps gradient propagation is not by itself evidence that the trained model generalizes better.
What a Strong Answer Covers Guidance
The role of the chain rule and repeated Jacobian products in shrinking earlier-layer gradients.
Saturating activations, initialization and scale as possible contributors.
Layer-wise gradient and activation diagnostics, plus checks for disconnected computation or implementation errors.
Architecture-appropriate mitigations and evaluation after changing the training setup.
Follow-up Questions Guidance
Why does clipping gradients usually address a different problem?
Can a rectified activation still leave a unit unable to learn?