Return the First Successful Result from Parallel API Calls Under a Global Timeout
Company: Apple
Role: Software Engineer
Category: Software Engineering Fundamentals
Difficulty: medium
Interview Round: Technical Screen
A service needs a value from a backend API whose calls are slow and sometimes fail with an error. To get an answer quickly and reliably, it makes several of these calls in parallel and uses the first one that succeeds.
Implement a function that takes a list of such calls and an overall timeout, and that:
- starts all of the calls concurrently;
- returns the value of the first call that completes without an error;
- enforces the timeout on the operation as a whole, not per call: once the timeout has elapsed, the function must return control to its caller even if some calls are still running;
- reports a failure if every call fails, or if the timeout elapses before any call succeeds.
You may use any language and concurrency model (threads, futures or promises, async/await). A Python-style signature for reference:
```python
def first_success(calls: list[Callable[[], T]], timeout_seconds: float) -> T:
...
```
Each element of `calls` is a zero-argument function that performs one backend call and either returns a value or raises an exception.
```hint Keep one deadline
Compute the deadline once, when the function starts, and derive every wait from the time that is left instead of giving each wait the full timeout.
```
```hint A failure is not an answer
Decide what your waiting logic does when the call that finishes first has failed while other calls are still running.
```
```hint The calls you no longer need
After you have a winner, some calls may still be running. Think about what happens to them, and whether anything in your function waits for them before returning.
```
### Constraints and Clarifications
- Each call may take far longer than the timeout, and some calls may never return at all.
- A call that raises an exception has failed; a call that returns normally has succeeded.
- The calls are independent of each other, and the function is expected to return its answer to the caller as soon as it is known.
### Clarifying Questions
- When every call fails, should the caller receive all of the individual errors (for example, aggregated in one exception), only the last one, or a generic failure?
- Should calls that are still running be cancelled once a winner is found, and can the underlying client actually be interrupted in the middle of a request?
- Do the calls have side effects, so that several of them succeeding would be a problem?
- Is there a limit on how many calls may run at once, or on how much extra load the backend can absorb?
- Can a successful call legitimately return `None` or an empty value?
- If two calls succeed at almost the same moment, does it matter which value is returned?
### What a Strong Answer Covers
- One overall deadline enforced across all waits, with the function returning on time even when calls hang
- Correct handling of early failures: keep waiting for the remaining calls, and report failure as soon as every call has failed instead of waiting out the timeout
- Failure reporting that distinguishes "every call failed" from "timed out" and preserves the underlying errors
- Clean-up of calls that are no longer needed: cancellation where the runtime allows it, no blocking on them at return, no leaked threads or tasks
- Correct use of the chosen concurrency primitives, including thread safety of any shared state
- Resource cost: workers per invocation and the extra backend load created by parallel calls
- A deterministic way to test the timing-dependent paths
### Follow-up Questions
- Instead of starting all calls at once, start the next call only if nothing has succeeded after a short delay (a hedged request). How does your implementation change, and what does this save?
- Only `k` of the `n` calls may be in flight at any time. How do you schedule the rest while still honoring the overall deadline?
- The backend call is not idempotent (for example, it creates a record). Is this pattern still safe, and what would you require from the API?
- How would you unit-test the timeout and all-calls-failed paths without real sleeps or a real backend?
Overview: Implement a function that runs several slow, failure-prone backend API calls concurrently and returns the first successful result within a single overall timeout. It tests concurrency primitives, deadline handling, reporting when every call fails, and cleaning up calls that are still in flight, using threads, futures, or async code.