Make a Charge-Then-Record Payment Endpoint Safe Against Timeouts, Retries and Crashes
Company: Mercor
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
A payment endpoint in a web application works in two steps. It first calls an external payment provider to charge the customer, and then writes a record of the payment to the application's own database. You are asked what is wrong with this design and how to fix it.
The flow, reconstructed as a sketch:
```python
def pay(order_id, user_id, amount):
charge = provider.charge(user_id, amount) # external HTTP call
db.insert_payment(order_id, user_id, amount, charge.id, status="paid")
return {"status": "paid"}
```
Assume the provider accepts an idempotency key on charge requests, lets you look up the status of a charge, and can notify you of charge results through webhooks. Confirm what the real provider offers before relying on any of these.
### Clarifying Questions
- Who retries a failed or timed-out request: the browser or mobile client, this server, or both?
- Can the provider report a charge as pending, or does every charge call end in success or a decline?
- For how long does the provider remember an idempotency key, and what does it return when it sees the same key again?
- Is each order paid by exactly one payment, so that an order must never be charged twice?
### Part 1 — What can go wrong
List the failure modes of the code above. For each one, say what state the customer's money and your database end up in.
```hint Walk the timeline
Mark every point where the process can crash or a network call can end without a clear answer, and ask what the provider and your database each believe at that point.
```
#### What This Part Should Cover
- Failures between the external charge and the local write, including the case where the outcome of the charge is unknown
- How retries and concurrent submissions of the same payment behave
- Whether a local database transaction can cover the external charge
### Part 2 — Make retries safe
Redesign the flow so that a client or server retry of the same payment never produces a second charge or a second record.
```hint Name the attempt
A retry only helps if every party can tell it is the same payment as before. Think about who should create that identity, and where each party stores it.
```
#### What This Part Should Cover
- Where the identity of a payment attempt comes from and how long it lives
- How the server detects a duplicate request, including two duplicates that arrive at the same moment
- How the call to the provider is itself protected against duplicates
### Part 3 — Never lose track of a charge
Some payments will end with an unknown outcome: the process crashed mid-flow, or the provider call timed out. Change the flow so that no charge attempt can be lost track of, and explain how each unknown outcome is eventually resolved.
```hint States, not booleans
Think about which states a payment can be in between "requested" and "done", and which events move it from one state to the next.
```
#### What This Part Should Cover
- The states of a payment record and the allowed transitions between them
- Every path by which an unknown outcome gets resolved, and who owns each path
- Handling the same result arriving more than once, or results arriving in an unexpected order
### Part 4 — Where the lock goes
Duplicate requests for the same payment, a webhook, and a background recovery process may all try to update the same record. Where would you put a lock, and should it be held until the external provider call finishes?
```hint Count what the lock pins
Consider what a lock or an open transaction holds on to (connections, rows, other waiting requests) while a slow provider call is in flight.
```
#### What This Part Should Cover
- What actually needs mutual exclusion, and at what granularity
- The cost of holding a lock or database transaction across a network call
- Ways to get the same safety without a long-held lock
### What a Strong Answer Covers
- A complete inventory of failure points that separates a definite failure from an unknown outcome
- One payment identity shared end to end by the client, the server's storage and the provider call
- A persisted payment state machine whose transitions are safe to apply more than once
- A recovery path for every non-final state, with an owner and a time bound
- Locking that keeps critical sections short and never spans the external call
- Observability for stuck and mismatched payments
### Follow-up Questions
- If creating the order fails after the charge has succeeded, what do you do?
- How would you test the timeout and crash paths before they happen in production?
- What metrics and alerts would tell you that payments are stuck, or that your records and the provider's disagree?
- How does the design change if the provider does not support idempotency keys at all?
Overview: A payment endpoint calls an external provider to charge the customer and then writes a record to its own database. The question asks what can go wrong and how to redesign the flow, testing failure analysis for crashes and timeouts, safe retries, durable payment state, recovery through status checks, webhooks and reconciliation, and lock placement.
Read the full Mercor Software Engineer interview experience this question came from