Explain your role in a production incident across detection, impact assessment, mitigation, diagnosis, recovery validation, communication, and follow-up prevention work.
## Question
Describe a production incident and your role in it. Explain detection, impact assessment, mitigation, diagnosis, recovery validation, communication, and the changes made afterward. Separate what you personally did from the incident team's work.
### Constraints & Assumptions
- Protect confidential details and avoid blaming an individual.
- Preserve a timeline and distinguish mitigation from root-cause correction.
- Include one decision made under uncertainty.
- Do not claim the incident was successful merely because service recovered.
### Clarifying Questions to Ask
- Would you like a reliability, data-integrity, security, or deployment incident?
- Should I emphasize technical diagnosis or incident coordination?
- How much timeline detail is useful?
```hint Tell the story in operational order
Detection, containment, diagnosis, recovery, and prevention answer different questions. Do not collapse them into one fix.
```
### What a Strong Answer Covers
- Clear impact and role assignment.
- Safe mitigation chosen with incomplete information.
- Evidence-based diagnosis and rejected hypotheses.
- Recovery checks that include user behavior and data integrity.
- Blameless corrective actions with owners and verification.
### Follow-up Questions
1. What was the riskiest mitigation option you rejected?
2. How did you communicate uncertainty?
3. Which alert should have fired earlier?
4. What corrective action had the highest leverage?
Quick Answer: Explain your role in a production incident across detection, impact assessment, mitigation, diagnosis, recovery validation, communication, and follow-up prevention work.
Describe a production incident and your role in it. Explain detection, impact assessment, mitigation, diagnosis, recovery validation, communication, and the changes made afterward. Separate what you personally did from the incident team's work.
Constraints & Assumptions
Protect confidential details and avoid blaming an individual.
Preserve a timeline and distinguish mitigation from root-cause correction.
Include one decision made under uncertainty.
Do not claim the incident was successful merely because service recovered.
Clarifying Questions to Ask Guidance
Would you like a reliability, data-integrity, security, or deployment incident?
Should I emphasize technical diagnosis or incident coordination?
How much timeline detail is useful?
What a Strong Answer Covers Guidance
Clear impact and role assignment.
Safe mitigation chosen with incomplete information.
Evidence-based diagnosis and rejected hypotheses.
Recovery checks that include user behavior and data integrity.
Blameless corrective actions with owners and verification.
Follow-up Questions Guidance
What was the riskiest mitigation option you rejected?