Design a Data Puller for a Third-Party API Capped at 3 Requests per Second

Read the full interview experience this question came from →

Quick Overview

Design a service that pulls and syncs data from a third-party API that enforces a hard limit of 3 requests per second. It tests global rate limiting across distributed workers, retry and backoff behavior, incremental sync with checkpoints, and prioritizing work when demand exceeds the request budget.

Design a Data Puller for a Third-Party API Capped at 3 Requests per Second

Company: Attentive

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design a service that pulls data from a third-party provider's API into your own system and keeps it up to date. The provider enforces a hard rate limit of 3 requests per second on its API. The interviewer kept the discussion focused on one question: how does your design handle that limit? ### Constraints and Clarifications - The limit is hard: requests above it are rejected by the provider, so it cannot be exceeded even briefly during bursts. - What data is pulled, how much of it, and how fresh it must be are not given. Ask, or state your assumptions. ### Clarifying Questions - Does the limit apply to the whole integration (one API key), or separately to each customer account or credential you pull on behalf of? - How is the limit measured: per calendar second, or over any sliding one-second interval? What does the provider return when you exceed it, and does it say when to retry? - How much data must be pulled, and how fresh must it be: a one-time backfill, a periodic sync, or near real time? - Does the API offer pagination with large page sizes, batch endpoints, filtering by last-updated time, or webhooks? - Who consumes the pulled data, and what happens downstream when it is late? ### Part 1 — The pull pipeline Design the components that decide what to fetch, call the provider, and store the results. ```hint Remember where each sync stopped Think about what you must record between runs so that each run fetches only what changed, and so that a crash halfway through does not force a restart from scratch. ``` #### What This Part Should Cover - Components (scheduler, work queue, fetch workers, storage) and their data model, including sync state and checkpoints - Full backfill versus incremental sync - Idempotent writes, so that fetching the same page twice is harmless - How downstream consumers learn about new data ### Part 2 — Staying under 3 requests per second Many workers run in parallel, possibly on several machines. How do you guarantee that together they never send more than 3 requests per second to the provider, and what do you do when a request is rejected anyway? ```hint One budget, many callers Per-worker limits stop adding up to a global guarantee as soon as the number of workers changes. Decide where the single source of truth for the budget lives. ``` ```hint Rejections will still happen Network jitter, retries and other clients sharing the same key can all push you over. Plan the response to a rejection so it does not make the overload worse. ``` #### Clarifying Questions for this Part - Do any other systems use the same API key, and therefore share the same limit? - Would the provider penalize repeated violations, for example by suspending the key? #### What This Part Should Cover - A global rate-limiting mechanism and how workers obtain permission to send - Pacing that matches how the provider measures the limit, with headroom - Handling rejections: retry hints, backoff with jitter, retries counted against the budget - Behavior when the rate limiter itself fails ### Part 3 — When demand exceeds the budget At 3 requests per second, the provider allows at most 10,800 requests per hour. What happens when the work you want to do needs more than that? ```hint Spend requests where they matter Compare the cost of a request with the value of the data it returns, and look for ways to get more data per request or to skip requests whose answer has not changed. ``` #### What This Part Should Cover - Arithmetic connecting the request budget to data volume and required freshness - Getting more out of each request, and avoiding requests that return nothing new - Prioritization and fairness, so that one large backfill does not starve fresh updates - Monitoring backlog and data lag, and what consumers are told when data falls behind ### What a Strong Answer Covers - Clarifies the scope and measurement of the limit before designing around it - A single enforceable global limit rather than per-process limits - Resilience: checkpoints, idempotent writes, and retries that respect the provider - Concrete numbers linking the request budget to data volume and freshness - Observability: sent rate, rejection rate, queue depth and data lag per tenant ### Follow-up Questions - Without warning, the provider starts rejecting requests whenever you exceed 2 per second. How does your system detect this and adapt? - A new customer needs a backfill that would take days at the current rate. How do you schedule it without hurting existing customers' freshness? - The central rate limiter is down. Do the workers stop, or continue with a local limit? What are the risks of each choice? - How would the design change if the limit were 3 requests per second per customer account instead of one global limit?

Overview: Design a service that pulls and syncs data from a third-party API that enforces a hard limit of 3 requests per second. It tests global rate limiting across distributed workers, retry and backoff behavior, incremental sync with checkpoints, and prioritizing work when demand exceeds the request budget.

Read the full Attentive Software Engineer interview experience this question came from

|Home/System Design/Attentive
Attentive logo
Attentive
Sep 7, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Design a service that pulls data from a third-party provider's API into your own system and keeps it up to date. The provider enforces a hard rate limit of 3 requests per second on its API. The interviewer kept the discussion focused on one question: how does your design handle that limit?

Constraints and Clarifications

  • The limit is hard: requests above it are rejected by the provider, so it cannot be exceeded even briefly during bursts.
  • What data is pulled, how much of it, and how fresh it must be are not given. Ask, or state your assumptions.

Clarifying Questions Guidance

  • Does the limit apply to the whole integration (one API key), or separately to each customer account or credential you pull on behalf of?
  • How is the limit measured: per calendar second, or over any sliding one-second interval? What does the provider return when you exceed it, and does it say when to retry?
  • How much data must be pulled, and how fresh must it be: a one-time backfill, a periodic sync, or near real time?
  • Does the API offer pagination with large page sizes, batch endpoints, filtering by last-updated time, or webhooks?
  • Who consumes the pulled data, and what happens downstream when it is late?

Part 1 — The pull pipeline

Design the components that decide what to fetch, call the provider, and store the results.

What This Part Should Cover Guidance

  • Components (scheduler, work queue, fetch workers, storage) and their data model, including sync state and checkpoints
  • Full backfill versus incremental sync
  • Idempotent writes, so that fetching the same page twice is harmless
  • How downstream consumers learn about new data

Part 2 — Staying under 3 requests per second

Many workers run in parallel, possibly on several machines. How do you guarantee that together they never send more than 3 requests per second to the provider, and what do you do when a request is rejected anyway?

Clarifying Questions for this Part Guidance

  • Do any other systems use the same API key, and therefore share the same limit?
  • Would the provider penalize repeated violations, for example by suspending the key?

What This Part Should Cover Guidance

  • A global rate-limiting mechanism and how workers obtain permission to send
  • Pacing that matches how the provider measures the limit, with headroom
  • Handling rejections: retry hints, backoff with jitter, retries counted against the budget
  • Behavior when the rate limiter itself fails

Part 3 — When demand exceeds the budget

At 3 requests per second, the provider allows at most 10,800 requests per hour. What happens when the work you want to do needs more than that?

What This Part Should Cover Guidance

  • Arithmetic connecting the request budget to data volume and required freshness
  • Getting more out of each request, and avoiding requests that return nothing new
  • Prioritization and fairness, so that one large backfill does not starve fresh updates
  • Monitoring backlog and data lag, and what consumers are told when data falls behind

What a Strong Answer Covers Guidance

  • Clarifies the scope and measurement of the limit before designing around it
  • A single enforceable global limit rather than per-process limits
  • Resilience: checkpoints, idempotent writes, and retries that respect the provider
  • Concrete numbers linking the request budget to data volume and freshness
  • Observability: sent rate, rejection rate, queue depth and data lag per tenant

Follow-up Questions Guidance

  • Without warning, the provider starts rejecting requests whenever you exceed 2 per second. How does your system detect this and adapt?
  • A new customer needs a backfill that would take days at the current rate. How do you schedule it without hurting existing customers' freshness?
  • The central rate limiter is down. Do the workers stop, or continue with a local limit? What are the risks of each choice?
  • How would the design change if the limit were 3 requests per second per customer account instead of one global limit?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...