Finding the Root Cause of an Intermittent Failure When the Obvious Metrics Look Healthy

Read the full interview experience this question came from →

Quick Overview

A behavioral debugging question: describe a system that failed intermittently while its obvious metrics looked healthy, how you found the root cause, and how long it took. It tests whether you can explain why the dashboards stayed green and show a systematic, evidence-driven narrowing process.

Finding the Root Cause of an Intermittent Failure When the Obvious Metrics Look Healthy

Company: Amazon

Role: Software Engineer

Category: Behavioral & Leadership

Difficulty: medium

Interview Round: Onsite

Tell me about a time when a system was failing intermittently, and the very obvious metrics were all looking good. How did you find the root cause, and how long did it take? This was one of three behavioral questions in an onsite round. When the first answer did not satisfy the interviewer, they restated the same question in different words and asked it again. Expect to be pushed until the debugging path itself is concrete: what you looked at, in what order, and what finally exposed the cause. ```hint Explain why the dashboards stayed green A strong story names the specific reason the obvious metrics missed the failure. Think about what the standard dashboards measure, where they measure it, and what they average away. ``` ```hint Show the narrowing Describe how you went from "something fails sometimes" to one component, one condition and one mechanism, including at least one hypothesis that turned out to be wrong. ``` ### Clarifying Questions - Does it have to be a system I owned, or can it be a dependency I debugged on behalf of my team? - Should the time I give include detecting the problem, or only finding the root cause once it was noticed? - How much technical detail do you want about the root cause itself? ### What a Strong Answer Covers - A concrete intermittent symptom and its impact on users or downstream systems - A specific explanation of why the obvious metrics looked healthy - A systematic narrowing process: slicing the data, forming and discarding hypotheses, reproducing the failure - A root cause explained down to the mechanism, with evidence that the fix removed the failure - An honest timeline, and a lasting change to monitoring or design so the next occurrence is caught sooner ### Follow-up Questions - Which signal finally exposed the failure, and why was no alarm watching it? - What was your first hypothesis, and what evidence ruled it out? - What made the investigation take as long as it did, and what would have shortened it? - What did you change afterward so that this class of failure appears on a dashboard or an alarm?

Overview: A behavioral debugging question: describe a system that failed intermittently while its obvious metrics looked healthy, how you found the root cause, and how long it took. It tests whether you can explain why the dashboards stayed green and show a systematic, evidence-driven narrowing process.

Read the full Amazon Software Engineer interview experience this question came from

|Home/Behavioral & Leadership/Amazon
Amazon logo
Amazon
Sep 10, 2026
mediumSoftware EngineerOnsiteBehavioral & Leadership
0
0

Tell me about a time when a system was failing intermittently, and the very obvious metrics were all looking good. How did you find the root cause, and how long did it take?

This was one of three behavioral questions in an onsite round. When the first answer did not satisfy the interviewer, they restated the same question in different words and asked it again. Expect to be pushed until the debugging path itself is concrete: what you looked at, in what order, and what finally exposed the cause.

Clarifying Questions Guidance

  • Does it have to be a system I owned, or can it be a dependency I debugged on behalf of my team?
  • Should the time I give include detecting the problem, or only finding the root cause once it was noticed?
  • How much technical detail do you want about the root cause itself?

What a Strong Answer Covers Guidance

  • A concrete intermittent symptom and its impact on users or downstream systems
  • A specific explanation of why the obvious metrics looked healthy
  • A systematic narrowing process: slicing the data, forming and discarding hypotheses, reproducing the failure
  • A root cause explained down to the mechanism, with evidence that the fix removed the failure
  • An honest timeline, and a lasting change to monitoring or design so the next occurrence is caught sooner

Follow-up Questions Guidance

  • Which signal finally exposed the failure, and why was no alarm watching it?
  • What was your first hypothesis, and what evidence ruled it out?
  • What made the investigation take as long as it did, and what would have shortened it?
  • What did you change afterward so that this class of failure appears on a dashboard or an alarm?
Loading comments...