Design a Data Puller for a Third-Party API Capped at 3 Requests per Second
Company: Attentive
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design a service that pulls data from a third-party provider's API into your own system and keeps it up to date. The provider enforces a hard rate limit of 3 requests per second on its API. The interviewer kept the discussion focused on one question: how does your design handle that limit?
### Constraints and Clarifications
- The limit is hard: requests above it are rejected by the provider, so it cannot be exceeded even briefly during bursts.
- What data is pulled, how much of it, and how fresh it must be are not given. Ask, or state your assumptions.
### Clarifying Questions
- Does the limit apply to the whole integration (one API key), or separately to each customer account or credential you pull on behalf of?
- How is the limit measured: per calendar second, or over any sliding one-second interval? What does the provider return when you exceed it, and does it say when to retry?
- How much data must be pulled, and how fresh must it be: a one-time backfill, a periodic sync, or near real time?
- Does the API offer pagination with large page sizes, batch endpoints, filtering by last-updated time, or webhooks?
- Who consumes the pulled data, and what happens downstream when it is late?
### Part 1 — The pull pipeline
Design the components that decide what to fetch, call the provider, and store the results.
```hint Remember where each sync stopped
Think about what you must record between runs so that each run fetches only what changed, and so that a crash halfway through does not force a restart from scratch.
```
#### What This Part Should Cover
- Components (scheduler, work queue, fetch workers, storage) and their data model, including sync state and checkpoints
- Full backfill versus incremental sync
- Idempotent writes, so that fetching the same page twice is harmless
- How downstream consumers learn about new data
### Part 2 — Staying under 3 requests per second
Many workers run in parallel, possibly on several machines. How do you guarantee that together they never send more than 3 requests per second to the provider, and what do you do when a request is rejected anyway?
```hint One budget, many callers
Per-worker limits stop adding up to a global guarantee as soon as the number of workers changes. Decide where the single source of truth for the budget lives.
```
```hint Rejections will still happen
Network jitter, retries and other clients sharing the same key can all push you over. Plan the response to a rejection so it does not make the overload worse.
```
#### Clarifying Questions for this Part
- Do any other systems use the same API key, and therefore share the same limit?
- Would the provider penalize repeated violations, for example by suspending the key?
#### What This Part Should Cover
- A global rate-limiting mechanism and how workers obtain permission to send
- Pacing that matches how the provider measures the limit, with headroom
- Handling rejections: retry hints, backoff with jitter, retries counted against the budget
- Behavior when the rate limiter itself fails
### Part 3 — When demand exceeds the budget
At 3 requests per second, the provider allows at most 10,800 requests per hour. What happens when the work you want to do needs more than that?
```hint Spend requests where they matter
Compare the cost of a request with the value of the data it returns, and look for ways to get more data per request or to skip requests whose answer has not changed.
```
#### What This Part Should Cover
- Arithmetic connecting the request budget to data volume and required freshness
- Getting more out of each request, and avoiding requests that return nothing new
- Prioritization and fairness, so that one large backfill does not starve fresh updates
- Monitoring backlog and data lag, and what consumers are told when data falls behind
### What a Strong Answer Covers
- Clarifies the scope and measurement of the limit before designing around it
- A single enforceable global limit rather than per-process limits
- Resilience: checkpoints, idempotent writes, and retries that respect the provider
- Concrete numbers linking the request budget to data volume and freshness
- Observability: sent rate, rejection rate, queue depth and data lag per tenant
### Follow-up Questions
- Without warning, the provider starts rejecting requests whenever you exceed 2 per second. How does your system detect this and adapt?
- A new customer needs a backfill that would take days at the current rate. How do you schedule it without hurting existing customers' freshness?
- The central rate limiter is down. Do the workers stop, or continue with a local limit? What are the risks of each choice?
- How would the design change if the limit were 3 requests per second per customer account instead of one global limit?
Overview: Design a service that pulls and syncs data from a third-party API that enforces a hard limit of 3 requests per second. It tests global rate limiting across distributed workers, retry and backoff behavior, incremental sync with checkpoints, and prioritizing work when demand exceeds the request budget.
Read the full Attentive Software Engineer interview experience this question came from