Meet a 10-second latency and 99% availability target across 3-5 dependent services

Read the full interview experience this question came from →

Quick Overview

An open-ended design discussion about an application that calls three to five dependent services and must keep latency under 10 seconds and availability at 99% or better. It tests objective definitions, availability and tail-latency math, timeouts, retries, fallbacks, and adapting as requirements change.

Meet a 10-second latency and 99% availability target across 3-5 dependent services

Company: Generalmotors

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

An application serves each request by calling three to five dependent services. How do you make the application meet specific service-level goals, for example end-to-end latency under 10 seconds and availability of at least 99%? This is less a standard "build a system" design round than an open-ended optimization discussion, held in a 45-minute slot. Once the goals are confirmed, the problem is broken down and discussed step by step. Expect the interviewer to switch to a new topic or add a requirement before the current thread is finished; keeping the changing requirements straight is part of the exercise. ### Constraints and Clarifications - The application depends on three to five services, the latency target is under 10 seconds, and the availability target is 99% or higher. Ask about any other numbers, or state them explicitly as assumptions. ### Clarifying Questions - Is the 10-second target a maximum or a percentile such as p99, and is it measured at the client or at the service? - Does 99% availability mean the fraction of successful requests or uptime over a calendar window, and over what period is it measured? - Which dependency calls are independent, and which need another call's output? - What latency distribution and availability does each dependency currently deliver, and does the team own any of them? - Are the requests reads, writes or both, and are the writes safe to retry? - Which dependencies are essential to a useful response, and which could be skipped or served stale? ### Part 1 — Latency under 10 seconds Explain how you turn the 10-second end-to-end target into a budget for each dependency call, and what you change in the application when a dependency is too slow to fit. ```hint Sequence versus fan-out Draw which calls depend on which. One path through that drawing sets the total latency, and percentiles do not add the way averages do. ``` #### What This Part Should Cover - The dependency graph and its critical path, with sequential and parallel calls distinguished - Per-call timeouts derived from the overall deadline, and propagation of that deadline downstream - Tail latency under fan-out, and techniques that cut it - What to do when the inherent work cannot finish within 10 seconds ### Part 2 — Availability of at least 99% Explain how the application reaches 99% availability when every request depends on several services, any of which can fail. ```hint Compose the dependencies Before choosing any mitigation, estimate what the whole chain delivers if every dependency is exactly as available as your own target. ``` #### What This Part Should Cover - How dependency availabilities combine, and the error budget that remains - Mitigations that keep a failed dependency from failing the request - Protecting the application from a slow or failing dependency - How both objectives are measured and alerted on ### What a Strong Answer Covers - Exact definitions of both objectives before any design is proposed - Quantitative reasoning: composed availability, tail-latency amplification, and the size of the error budget - Mitigations matched to which dependencies are critical and which calls are safe to retry - How the two goals interact, since retries and fallbacks spend latency budget - A running summary of requirements and decisions, against which each new requirement is checked ### Follow-up Questions - A sixth dependency is added to the critical path. What happens to both targets, and what do you change? - One dependency's p99 latency doubles overnight. How do you find out, and what protects the 10-second target in the meantime? - How would you decide the latency and availability objectives to require from each dependency's owning team?

Overview: An open-ended design discussion about an application that calls three to five dependent services and must keep latency under 10 seconds and availability at 99% or better. It tests objective definitions, availability and tail-latency math, timeouts, retries, fallbacks, and adapting as requirements change.

Read the full Generalmotors Software Engineer interview experience this question came from

|Home/System Design/Generalmotors
Generalmotors logo
Generalmotors
Sep 9, 2026
mediumSoftware EngineerOnsiteSystem Design
1
0

An application serves each request by calling three to five dependent services. How do you make the application meet specific service-level goals, for example end-to-end latency under 10 seconds and availability of at least 99%?

This is less a standard "build a system" design round than an open-ended optimization discussion, held in a 45-minute slot. Once the goals are confirmed, the problem is broken down and discussed step by step. Expect the interviewer to switch to a new topic or add a requirement before the current thread is finished; keeping the changing requirements straight is part of the exercise.

Constraints and Clarifications

  • The application depends on three to five services, the latency target is under 10 seconds, and the availability target is 99% or higher. Ask about any other numbers, or state them explicitly as assumptions.

Clarifying Questions Guidance

  • Is the 10-second target a maximum or a percentile such as p99, and is it measured at the client or at the service?
  • Does 99% availability mean the fraction of successful requests or uptime over a calendar window, and over what period is it measured?
  • Which dependency calls are independent, and which need another call's output?
  • What latency distribution and availability does each dependency currently deliver, and does the team own any of them?
  • Are the requests reads, writes or both, and are the writes safe to retry?
  • Which dependencies are essential to a useful response, and which could be skipped or served stale?

Part 1 — Latency under 10 seconds

Explain how you turn the 10-second end-to-end target into a budget for each dependency call, and what you change in the application when a dependency is too slow to fit.

What This Part Should Cover Guidance

  • The dependency graph and its critical path, with sequential and parallel calls distinguished
  • Per-call timeouts derived from the overall deadline, and propagation of that deadline downstream
  • Tail latency under fan-out, and techniques that cut it
  • What to do when the inherent work cannot finish within 10 seconds

Part 2 — Availability of at least 99%

Explain how the application reaches 99% availability when every request depends on several services, any of which can fail.

What This Part Should Cover Guidance

  • How dependency availabilities combine, and the error budget that remains
  • Mitigations that keep a failed dependency from failing the request
  • Protecting the application from a slow or failing dependency
  • How both objectives are measured and alerted on

What a Strong Answer Covers Guidance

  • Exact definitions of both objectives before any design is proposed
  • Quantitative reasoning: composed availability, tail-latency amplification, and the size of the error budget
  • Mitigations matched to which dependencies are critical and which calls are safe to retry
  • How the two goals interact, since retries and fallbacks spend latency budget
  • A running summary of requirements and decisions, against which each new requirement is checked

Follow-up Questions Guidance

  • A sixth dependency is added to the critical path. What happens to both targets, and what do you change?
  • One dependency's p99 latency doubles overnight. How do you find out, and what protects the 10-second target in the meantime?
  • How would you decide the latency and availability objectives to require from each dependency's owning team?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...