Practice triaging a customer-reported cloud connection incident as an on-call PM. The solution covers intake, severity, timestamps, regions, endpoints, network and auth hypotheses, diagnostics, logs, metrics, mitigation, internal and external communication, workarounds, decision logs, and postmortem follow-up.
##### Question
Walk through how you would triage a customer-reported cloud connection issue. Cover the initial information you need, hypotheses you would test, diagnostic tools or logs you would inspect, and how you would communicate status to stakeholders.
Quick Answer: Practice triaging a customer-reported cloud connection incident as an on-call PM. The solution covers intake, severity, timestamps, regions, endpoints, network and auth hypotheses, diagnostics, logs, metrics, mitigation, internal and external communication, workarounds, decision logs, and postmortem follow-up.
List hypotheses and diagnostic tools, metrics, and logs to inspect.
What This Part Should Cover Guidance
Client/network issues, DNS, TLS, firewall/proxy, VPN/PrivateLink, auth/permissions, rate limits, service errors, dependency outages, deploy/config regression, and regional incidents.
Metrics such as error rate, latency, connection failures, saturation, request volume, and success by region.
Logs, traces, request IDs, load balancer metrics, network telemetry, deployment timeline, and status pages.
Part 3 - Communication and Follow-Up
Explain how you communicate status and next steps internally and externally.
What This Part Should Cover Guidance
Incident channel, roles, cadence, customer updates, support macros, executive summary, and escalation.
Workaround communication.
Decision log and timeline.
Post-incident review, root cause, corrective actions, and prevention.
What a Strong Answer Covers Guidance
A strong answer scopes impact fast, tests layered hypotheses, drives mitigation, keeps stakeholders informed, and turns the incident into follow-up actions that reduce future recurrence.
Follow-up Questions Guidance
What would you do if only one enterprise customer is affected?
What if Support reports timeouts but SRE dashboards look normal?
How would you decide whether to page another team?
What do you say externally before root cause is known?