Meet a 10-second latency and 99% availability target across 3-5 dependent services
Company: Generalmotors
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
An application serves each request by calling three to five dependent services. How do you make the application meet specific service-level goals, for example end-to-end latency under 10 seconds and availability of at least 99%?
This is less a standard "build a system" design round than an open-ended optimization discussion, held in a 45-minute slot. Once the goals are confirmed, the problem is broken down and discussed step by step. Expect the interviewer to switch to a new topic or add a requirement before the current thread is finished; keeping the changing requirements straight is part of the exercise.
### Constraints and Clarifications
- The application depends on three to five services, the latency target is under 10 seconds, and the availability target is 99% or higher. Ask about any other numbers, or state them explicitly as assumptions.
### Clarifying Questions
- Is the 10-second target a maximum or a percentile such as p99, and is it measured at the client or at the service?
- Does 99% availability mean the fraction of successful requests or uptime over a calendar window, and over what period is it measured?
- Which dependency calls are independent, and which need another call's output?
- What latency distribution and availability does each dependency currently deliver, and does the team own any of them?
- Are the requests reads, writes or both, and are the writes safe to retry?
- Which dependencies are essential to a useful response, and which could be skipped or served stale?
### Part 1 — Latency under 10 seconds
Explain how you turn the 10-second end-to-end target into a budget for each dependency call, and what you change in the application when a dependency is too slow to fit.
```hint Sequence versus fan-out
Draw which calls depend on which. One path through that drawing sets the total latency, and percentiles do not add the way averages do.
```
#### What This Part Should Cover
- The dependency graph and its critical path, with sequential and parallel calls distinguished
- Per-call timeouts derived from the overall deadline, and propagation of that deadline downstream
- Tail latency under fan-out, and techniques that cut it
- What to do when the inherent work cannot finish within 10 seconds
### Part 2 — Availability of at least 99%
Explain how the application reaches 99% availability when every request depends on several services, any of which can fail.
```hint Compose the dependencies
Before choosing any mitigation, estimate what the whole chain delivers if every dependency is exactly as available as your own target.
```
#### What This Part Should Cover
- How dependency availabilities combine, and the error budget that remains
- Mitigations that keep a failed dependency from failing the request
- Protecting the application from a slow or failing dependency
- How both objectives are measured and alerted on
### What a Strong Answer Covers
- Exact definitions of both objectives before any design is proposed
- Quantitative reasoning: composed availability, tail-latency amplification, and the size of the error budget
- Mitigations matched to which dependencies are critical and which calls are safe to retry
- How the two goals interact, since retries and fallbacks spend latency budget
- A running summary of requirements and decisions, against which each new requirement is checked
### Follow-up Questions
- A sixth dependency is added to the critical path. What happens to both targets, and what do you change?
- One dependency's p99 latency doubles overnight. How do you find out, and what protects the 10-second target in the meantime?
- How would you decide the latency and availability objectives to require from each dependency's owning team?
Overview: An open-ended design discussion about an application that calls three to five dependent services and must keep latency under 10 seconds and availability at 99% or better. It tests objective definitions, availability and tail-latency math, timeouts, retries, fallbacks, and adapting as requirements change.
Read the full Generalmotors Software Engineer interview experience this question came from