Design an API to Deep-Copy a Dataset with Its Tasks, Model Responses and Grades

Read the full interview experience this question came from →

Quick Overview

A system design question about an API that copies a dataset together with its tasks, model responses and grades so that the copy can be edited independently. It covers ID remapping, background jobs with progress tracking and resumable checkpoints, idempotent requests, worker leases, edits to the source during the copy, and cleanup after failure.

Design an API to Deep-Copy a Dataset with Its Tasks, Model Responses and Grades

Company: Mercor

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Your application stores datasets. Each dataset contains tasks, each task has model responses, and each response has grades. Design an API that copies a dataset together with all of its tasks, model responses and grades. After the copy, the new dataset and the original must be independently editable: a change to one must never show up in the other. Assume the data lives in a relational database and every row references its parent by ID. ### Clarifying Questions - Is the hierarchy strictly dataset, then tasks, then responses, then grades? Do any rows also reference entities (graders, rubrics, users) that should be shared rather than copied? - How large can a dataset get, and how long may a copy take before the caller needs an answer? - Must the copy reflect a single point in time of the source, or is a copy that picks up some edits made during the copy acceptable? - Do rows reference large payloads stored outside the database, such as files or long texts? - Who may copy a dataset, and who owns the copy? ### Part 1 — The copy API and ID remapping Define the endpoint and the copy procedure for a dataset small enough to copy within one request. Every copied row gets a new ID, and every reference inside the copy must point to the copied parent, not to the original. ```hint Order matters A copied response has to point at a copied task, and that task did not exist a moment ago. Think about the order in which rows are copied and what you need to remember about each one. ``` #### What This Part Should Cover - The request and response contract of the copy endpoint - The order of copying, and how references are rewritten to new IDs - Atomicity in the small case, so that nobody sees a half-copied dataset ### Part 2 — Large datasets Some datasets are too large to copy within one request or one database transaction. Change the design so the copy runs in the background, the caller can track progress, and a crash halfway through does not mean starting over. ```hint Resume from where? After a crash, the job must know exactly which rows were already copied and what their new IDs are. Think about what must be durable, and in which transaction it is written. ``` #### What This Part Should Cover - The job lifecycle, and the API to start it and check on it - Batching, and a durable checkpoint that stays consistent with the rows already copied - How the old-to-new ID mapping survives a restart ### Part 3 — Duplicate requests and competing workers The client may send the same copy request twice (a retry or a double click), and several background workers pick up jobs. Make sure one request creates only one copy, and that two workers never run the same job at the same time. ```hint Locks that outlive their holder A worker can crash or stall while it holds a job. Think about how the others find out, and what happens if the stalled worker wakes up and keeps writing. ``` #### What This Part Should Cover - Deduplicating requests, and what the second request receives - How a worker claims a job, keeps it, and loses it - Preventing a stale worker from writing after its claim has passed to another worker ### Part 4 — Edits during the copy, and failures The source dataset may be edited while the copy is running. What does the copy contain? And when a job fails, do you keep the partially copied rows and continue later, or delete them? ```hint Define the snapshot Decide what "a copy of the dataset" means while the dataset is changing, then check which rows your batching would pick up or miss. ``` #### What This Part Should Cover - The consistency the copy promises, and how it is enforced - What users can see of a copy that is still running or has failed - When to resume and when to clean up, and how cleanup itself is made safe ### What a Strong Answer Covers - A clear contract for copy semantics (deep copy, new IDs, independence) and an API to start and track a copy - A data model for jobs and for the ID mapping that makes progress durable and resumable - Idempotency at the request level and exclusive execution at the worker level, including protection against stale workers - An explicit consistency choice for edits made during the copy, with its cost - Failure handling that never exposes a half-copied dataset - Progress reporting and operational metrics ### Follow-up Questions - How would you make copies cheaper if most copies are never edited afterwards? - How would you cancel a running copy, and what happens to the rows already written? - If tasks, responses and grades lived in different services or database shards, what would change in the mapping and the copy order?

Overview: A system design question about an API that copies a dataset together with its tasks, model responses and grades so that the copy can be edited independently. It covers ID remapping, background jobs with progress tracking and resumable checkpoints, idempotent requests, worker leases, edits to the source during the copy, and cleanup after failure.

Read the full Mercor Software Engineer interview experience this question came from

|Home/System Design/Mercor
Mercor logo
Mercor
May 31, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Your application stores datasets. Each dataset contains tasks, each task has model responses, and each response has grades. Design an API that copies a dataset together with all of its tasks, model responses and grades. After the copy, the new dataset and the original must be independently editable: a change to one must never show up in the other.

Assume the data lives in a relational database and every row references its parent by ID.

Clarifying Questions Guidance

  • Is the hierarchy strictly dataset, then tasks, then responses, then grades? Do any rows also reference entities (graders, rubrics, users) that should be shared rather than copied?
  • How large can a dataset get, and how long may a copy take before the caller needs an answer?
  • Must the copy reflect a single point in time of the source, or is a copy that picks up some edits made during the copy acceptable?
  • Do rows reference large payloads stored outside the database, such as files or long texts?
  • Who may copy a dataset, and who owns the copy?

Part 1 — The copy API and ID remapping

Define the endpoint and the copy procedure for a dataset small enough to copy within one request. Every copied row gets a new ID, and every reference inside the copy must point to the copied parent, not to the original.

What This Part Should Cover Guidance

  • The request and response contract of the copy endpoint
  • The order of copying, and how references are rewritten to new IDs
  • Atomicity in the small case, so that nobody sees a half-copied dataset

Part 2 — Large datasets

Some datasets are too large to copy within one request or one database transaction. Change the design so the copy runs in the background, the caller can track progress, and a crash halfway through does not mean starting over.

What This Part Should Cover Guidance

  • The job lifecycle, and the API to start it and check on it
  • Batching, and a durable checkpoint that stays consistent with the rows already copied
  • How the old-to-new ID mapping survives a restart

Part 3 — Duplicate requests and competing workers

The client may send the same copy request twice (a retry or a double click), and several background workers pick up jobs. Make sure one request creates only one copy, and that two workers never run the same job at the same time.

What This Part Should Cover Guidance

  • Deduplicating requests, and what the second request receives
  • How a worker claims a job, keeps it, and loses it
  • Preventing a stale worker from writing after its claim has passed to another worker

Part 4 — Edits during the copy, and failures

The source dataset may be edited while the copy is running. What does the copy contain? And when a job fails, do you keep the partially copied rows and continue later, or delete them?

What This Part Should Cover Guidance

  • The consistency the copy promises, and how it is enforced
  • What users can see of a copy that is still running or has failed
  • When to resume and when to clean up, and how cleanup itself is made safe

What a Strong Answer Covers Guidance

  • A clear contract for copy semantics (deep copy, new IDs, independence) and an API to start and track a copy
  • A data model for jobs and for the ID mapping that makes progress durable and resumable
  • Idempotency at the request level and exclusive execution at the worker level, including protection against stale workers
  • An explicit consistency choice for edits made during the copy, with its cost
  • Failure handling that never exposes a half-copied dataset
  • Progress reporting and operational metrics

Follow-up Questions Guidance

  • How would you make copies cheaper if most copies are never edited afterwards?
  • How would you cancel a running copy, and what happens to the rows already written?
  • If tasks, responses and grades lived in different services or database shards, what would change in the mapping and the copy order?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...