Writing up an earlier Mercor fullstack round. It was online, about half an hour, and after I passed I got an onsite. It was mostly reading code to find problems and talking about design. I'm putting this together from my notes, so the order may not match what actually happened.
- Payment API
Code that first calls an external provider to charge the customer, then writes a record to our own database. What's wrong with it?
For example, the charge succeeds but the DB write fails, or the provider times out and you don't know whether the charge went through, so would a retry charge twice? Our own DB transaction also can't roll back the external charge.
Ideas: the client generates an idempotency key for the same payment request and reuses it on retries, the server dedupes on it when persisting, and a fixed idempotency key is also used when calling the provider. Another approach is to store the payment record and its status first, then call the provider and update the record once you have the result. If something fails midway or the result is uncertain, how do you check the status, and how do you fill the gap through a webhook or reconciliation. Also where to put the lock, and whether to hold it all the way until the external call finishes.
- Multithreaded cache
An expensive function wrapped with a dict, roughly like this:
cache = {}
def compute(x):
if x in cache:
return cache[x]
result = expensive(x)
cache[x] = result
return result
What goes wrong when multiple threads call it at the same time, and how do you fix it? It leads into the same key being computed more than once, how to add a lock, and the scope of the lock.
- Spot the problems in CSV code
Given a piece of Python, roughly like this:
def to_csv(table):
output = ""
for row in table:
for cell in row:
output += cell
output += ","
output += "\n"
return output
"Spot issues with this code." That covers the extra trailing comma, the string concatenation (strings are immutable in Python), non-string input, and what to do when the content itself contains commas, quotes or newlines.
- Copying a dataset
Design an API that copies a dataset together with its tasks, model responses and grades, so that after the copy it can be modified independently.
It involves a mapping between old and new IDs: copy the tasks first, then the responses and grades, and swap the references to the new IDs. If the data is large, make it an async job that returns a job ID so you can check progress, copy in batches, and save the mapping and a checkpoint so it can resume if it dies midway.
There's also how to dedupe duplicate requests with an idempotency key, and how multiple workers avoid running the same job at the same time through a lease or lock. And what to do if the source data is modified during the copy, and whether a failed copy should keep the half-finished result and continue, or clean it up.
Discussion
Loading comments…