Finding the Root Cause of an Intermittent Failure When the Obvious Metrics Look Healthy
Company: Amazon
Role: Software Engineer
Category: Behavioral & Leadership
Difficulty: medium
Interview Round: Onsite
Tell me about a time when a system was failing intermittently, and the very obvious metrics were all looking good. How did you find the root cause, and how long did it take?
This was one of three behavioral questions in an onsite round. When the first answer did not satisfy the interviewer, they restated the same question in different words and asked it again. Expect to be pushed until the debugging path itself is concrete: what you looked at, in what order, and what finally exposed the cause.
```hint Explain why the dashboards stayed green
A strong story names the specific reason the obvious metrics missed the failure. Think about what the standard dashboards measure, where they measure it, and what they average away.
```
```hint Show the narrowing
Describe how you went from "something fails sometimes" to one component, one condition and one mechanism, including at least one hypothesis that turned out to be wrong.
```
### Clarifying Questions
- Does it have to be a system I owned, or can it be a dependency I debugged on behalf of my team?
- Should the time I give include detecting the problem, or only finding the root cause once it was noticed?
- How much technical detail do you want about the root cause itself?
### What a Strong Answer Covers
- A concrete intermittent symptom and its impact on users or downstream systems
- A specific explanation of why the obvious metrics looked healthy
- A systematic narrowing process: slicing the data, forming and discarding hypotheses, reproducing the failure
- A root cause explained down to the mechanism, with evidence that the fix removed the failure
- An honest timeline, and a lasting change to monitoring or design so the next occurrence is caught sooner
### Follow-up Questions
- Which signal finally exposed the failure, and why was no alarm watching it?
- What was your first hypothesis, and what evidence ruled it out?
- What made the investigation take as long as it did, and what would have shortened it?
- What did you change afterward so that this class of failure appears on a dashboard or an alarm?
Overview: A behavioral debugging question: describe a system that failed intermittently while its obvious metrics looked healthy, how you found the root cause, and how long it took. It tests whether you can explain why the dashboards stayed green and show a systematic, evidence-driven narrowing process.
Read the full Amazon Software Engineer interview experience this question came from