Triage a Suddenly Slow Production Service

Quick Overview

Triage a production service whose latency suddenly increased by scoping impact and comparing healthy and unhealthy dimensions. The solution connects traces, queues, dependencies, resources, recent changes, reversible mitigation, evidence preservation, recovery checks, and incident follow-up.

Triage a Suddenly Slow Production Service

Company: ByteDance

Role: Site Reliability Engineer

Category: Software Engineering Fundamentals

Difficulty: easy

Interview Round: Technical Screen

# Triage a Suddenly Slow Production Service A production service that normally meets its latency target has become slow. Describe how you would triage the incident from the first alert through mitigation and follow-up. You do not yet know whether the cause is application code, a downstream dependency, resource exhaustion, traffic shape, network behavior, or a recent change. Explain which evidence you gather first, how you narrow the scope, and when you mitigate before establishing the full root cause. ### Constraints & Assumptions - The service is still responding, but user-visible latency has materially increased. - Dashboards, logs, traces, deployment history, and host or container metrics are available but may be incomplete. - Changes made during the incident must be reversible and recorded. - Protecting users and preserving evidence are both incident goals. ### Clarifying Questions to Ask - Which latency percentile and endpoints are slow, and when did the change begin? - Is the problem global or limited by region, tenant, host, version, or request type? - Did error rate, saturation, queue depth, or throughput change at the same time? - What deployments, configuration changes, traffic events, or dependency incidents overlap the onset? - Which rollback, traffic-shift, or load-shedding controls are known to be safe? ### What a Strong Answer Covers - Confirmation of user impact and scope before drawing a cause from one aggregate graph. - Correlation of latency with traffic, errors, saturation, queues, resource pressure, and recent changes. - Comparison across healthy and unhealthy dimensions such as host, region, version, endpoint, or dependency. - Traces and structured logs used to divide time among queueing, application work, storage, and remote calls. - Safe mitigation choices including rollback, traffic shift, concurrency limits, cache use, or load shedding, with explicit risks. - Preservation of timestamps, commands, changes, and evidence while avoiding unbounded ad hoc production experiments. - Verification after mitigation and a follow-up plan for root cause, detection, capacity, and recurrence prevention. ### Follow-up Questions 1. CPU is low but latency and request queue depth are rising. What hypotheses become more likely? 2. Only one deployment version is slow, but rollback would discard a security fix. How do you reduce impact safely? 3. Database latency is elevated for this service but normal globally. What comparisons would you make next? 4. How do you distinguish a dependency timeout from local thread-pool starvation when both appear as slow requests? 5. Which evidence should be captured before restarting an unhealthy process?

Overview: Triage a production service whose latency suddenly increased by scoping impact and comparing healthy and unhealthy dimensions. The solution connects traces, queues, dependencies, resources, recent changes, reversible mitigation, evidence preservation, recovery checks, and incident follow-up.

|Home/Software Engineering Fundamentals/ByteDance
ByteDance logo
ByteDance
Sep 1, 2026
easySite Reliability EngineerTechnical ScreenSoftware Engineering Fundamentals
0
0

Triage a Suddenly Slow Production Service

A production service that normally meets its latency target has become slow. Describe how you would triage the incident from the first alert through mitigation and follow-up.

You do not yet know whether the cause is application code, a downstream dependency, resource exhaustion, traffic shape, network behavior, or a recent change. Explain which evidence you gather first, how you narrow the scope, and when you mitigate before establishing the full root cause.

Constraints & Assumptions

  • The service is still responding, but user-visible latency has materially increased.
  • Dashboards, logs, traces, deployment history, and host or container metrics are available but may be incomplete.
  • Changes made during the incident must be reversible and recorded.
  • Protecting users and preserving evidence are both incident goals.

Clarifying Questions to Ask Guidance

  • Which latency percentile and endpoints are slow, and when did the change begin?
  • Is the problem global or limited by region, tenant, host, version, or request type?
  • Did error rate, saturation, queue depth, or throughput change at the same time?
  • What deployments, configuration changes, traffic events, or dependency incidents overlap the onset?
  • Which rollback, traffic-shift, or load-shedding controls are known to be safe?

What a Strong Answer Covers Guidance

  • Confirmation of user impact and scope before drawing a cause from one aggregate graph.
  • Correlation of latency with traffic, errors, saturation, queues, resource pressure, and recent changes.
  • Comparison across healthy and unhealthy dimensions such as host, region, version, endpoint, or dependency.
  • Traces and structured logs used to divide time among queueing, application work, storage, and remote calls.
  • Safe mitigation choices including rollback, traffic shift, concurrency limits, cache use, or load shedding, with explicit risks.
  • Preservation of timestamps, commands, changes, and evidence while avoiding unbounded ad hoc production experiments.
  • Verification after mitigation and a follow-up plan for root cause, detection, capacity, and recurrence prevention.

Follow-up Questions Guidance

  1. CPU is low but latency and request queue depth are rising. What hypotheses become more likely?
  2. Only one deployment version is slow, but rollback would discard a security fix. How do you reduce impact safely?
  3. Database latency is elevated for this service but normal globally. What comparisons would you make next?
  4. How do you distinguish a dependency timeout from local thread-pool starvation when both appear as slow requests?
  5. Which evidence should be captured before restarting an unhealthy process?
Loading comments...