Explain Vanishing Gradients in Neural-Network Training

Quick Overview

Diagnose vanishing gradients through the chain rule, activation and gradient checks, and architecture-appropriate training remedies.

Explain Vanishing Gradients in Neural-Network Training

Company: Snapchat

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: hard

Interview Round: Technical Screen

# Explain Vanishing Gradients in Neural-Network Training What is the vanishing-gradient problem, why can it occur when training a neural network, and what would you do to diagnose and reduce it? Explain how gradients propagate through repeated layers or time steps and distinguish very small backpropagated gradients from a model that simply has a low loss near a useful optimum. You may illustrate the mechanism with a deep feed-forward network or a recurrent network, but state which setting you are using. No architecture or activation function is specified, so explain which proposed remedies depend on those choices. Discuss why a change that helps gradient propagation is not by itself evidence that the trained model generalizes better. ### What a Strong Answer Covers - The role of the chain rule and repeated Jacobian products in shrinking earlier-layer gradients. - Saturating activations, initialization and scale as possible contributors. - Layer-wise gradient and activation diagnostics, plus checks for disconnected computation or implementation errors. - Architecture-appropriate mitigations and evaluation after changing the training setup. ### Follow-up Questions - Why does clipping gradients usually address a different problem? - Can a rectified activation still leave a unit unable to learn?

Overview: Diagnose vanishing gradients through the chain rule, activation and gradient checks, and architecture-appropriate training remedies.

|Home/Machine Learning/Snapchat
Snapchat logo
Snapchat
Sep 12, 2026
hardMachine Learning EngineerTechnical ScreenMachine Learning
0
0

Explain Vanishing Gradients in Neural-Network Training

What is the vanishing-gradient problem, why can it occur when training a neural network, and what would you do to diagnose and reduce it? Explain how gradients propagate through repeated layers or time steps and distinguish very small backpropagated gradients from a model that simply has a low loss near a useful optimum.

You may illustrate the mechanism with a deep feed-forward network or a recurrent network, but state which setting you are using. No architecture or activation function is specified, so explain which proposed remedies depend on those choices. Discuss why a change that helps gradient propagation is not by itself evidence that the trained model generalizes better.

What a Strong Answer Covers Guidance

  • The role of the chain rule and repeated Jacobian products in shrinking earlier-layer gradients.
  • Saturating activations, initialization and scale as possible contributors.
  • Layer-wise gradient and activation diagnostics, plus checks for disconnected computation or implementation errors.
  • Architecture-appropriate mitigations and evaluation after changing the training setup.

Follow-up Questions Guidance

  • Why does clipping gradients usually address a different problem?
  • Can a rectified activation still leave a unit unable to learn?
Loading comments...