Identify What Does Not Fix Vanishing Gradients

Read the full interview experience this question came from →

Quick Overview

Explain why raising the learning rate does not fix vanishing gradients and compare skip connections, ReLU, and batch normalization.

Identify What Does Not Fix Vanishing Gradients

Company: C3 AI

Role: Data Scientist

Category: Machine Learning

Difficulty: medium

Interview Round: Online Assessment

# Identify What Does Not Fix Vanishing Gradients Which change is not a reliable remedy for vanishing gradients in a deep neural network: adding skip connections, using an activation such as ReLU, raising the learning rate, or adding batch normalization? Explain the mechanism behind your choice and the limitations of the other interventions. ### What a Strong Answer Covers - Gradient propagation through a product of layer derivatives. - Why increasing the optimizer step size does not repair a vanishing gradient signal. - How skip connections, activation choice, and normalization can help without guaranteeing a cure. ```hint Separate signal from step size The gradient is computed before the optimizer multiplies it by the learning rate. ``` ### Follow-up Questions - Can ReLU units still stop transmitting gradients? - Why can residual connections help very deep models?

Overview: Explain why raising the learning rate does not fix vanishing gradients and compare skip connections, ReLU, and batch normalization.

Read the full C3 AI Data Scientist interview experience this question came from

|Home/Machine Learning/C3 AI
C3 AI logo
C3 AI
Sep 15, 2026
mediumData ScientistOnline AssessmentMachine Learning
0
0

Identify What Does Not Fix Vanishing Gradients

Which change is not a reliable remedy for vanishing gradients in a deep neural network: adding skip connections, using an activation such as ReLU, raising the learning rate, or adding batch normalization? Explain the mechanism behind your choice and the limitations of the other interventions.

What a Strong Answer Covers Guidance

  • Gradient propagation through a product of layer derivatives.
  • Why increasing the optimizer step size does not repair a vanishing gradient signal.
  • How skip connections, activation choice, and normalization can help without guaranteeing a cure.

Follow-up Questions Guidance

  • Can ReLU units still stop transmitting gradients?
  • Why can residual connections help very deep models?
Loading comments...