Paying Dashers Reliably When the Payment Service Is Unavailable
Company: DoorDash
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
A delivery platform computes each Dasher's (delivery driver's) pay from their order events: pay accrues per minute in proportion to the number of orders in progress, and peak-time minutes pay double. Once an amount is computed, the platform asks a separate payment service to actually pay the Dasher.
After the pay calculation is coded, the interviewer asks: what if the payment service is unavailable? Explain what can go wrong, and design the hand-off from pay calculation to payout so that every Dasher is eventually paid exactly the right amount, once, even while the payment service is down or unreliable.
```hint Failure is not binary
Consider a call that times out: you do not know whether the money moved. Think about what the payment service must support for a retry to be safe.
```
```hint Decouple the steps
Ask what must be recorded durably on your side before you call the payment service, so that nothing is lost if your own process crashes during an outage.
```
### Clarifying Questions
- When are Dashers paid: per order, once a day, weekly, or on demand?
- Does the payment service accept an idempotency key, and can a payout's status be looked up by that key?
- How long may a payout be delayed before it becomes a problem for Dashers, and what should the Dasher app show in the meantime?
- Is the payment service internal, or a third-party provider with its own rate limits?
### What a Strong Answer Covers
- Distinguishing outright failures, timeouts with an unknown outcome, and slow responses
- A durable record of what is owed, written before any payout call
- Retries that cannot double-pay, and protection against overloading the service when it recovers
- Reconciliation between the platform's records and the payment service's records
- What the Dasher sees, and the alerting that tells operators when payouts are stuck
### Follow-up Questions
- The payment service comes back after several hours with a large backlog. How do you drain it without overwhelming the service?
- A bug in the pay calculation is found after some payouts have gone out. How do you correct underpayments and overpayments?
- Could you switch to a second payment provider during an outage, and what new risks would that bring?
Overview: A system design follow-up about paying delivery drivers when the downstream payment service is unavailable. It tests reasoning about timeouts with unknown outcomes, durable records of amounts owed, retries that must never double-pay, recovery after long outages, reconciliation, and what drivers and operators see while payouts are delayed.