Design a Data Processing System for JSON User Records: Batch vs Streaming Input

Read the full interview experience this question came from →

Quick Overview

Design a data processing system for JSON user records that contain names, addresses and payment methods, and explain how the design changes when the upstream source delivers data in batches versus as a continuous stream. It tests pipeline architecture, idempotent loading, ordering and late data, and the handling of sensitive fields.

Design a Data Processing System for JSON User Records: Batch vs Streaming Input

Company: Patreon

Role: Data Engineer

Category: System Design

Difficulty: easy

Interview Round: Technical Screen

Your team receives user records as JSON. Each record holds personal data such as a name, a mailing address and one or more payment methods. The round opened with a short exercise in which the candidate processed a large JSON dump of such records using plain Python loops. The interviewer then asked a brief, high-level design question: How would you build a data processing system for this data, and how would you build it if the upstream source delivers the data in batches, versus as a stream? A short, structured answer was expected rather than a deep dive. The record below only illustrates the shape of the data; the exact fields were not reported. ```json { "id": "u-001", "name": "Sample User", "address": {"street": "1 Main St", "city": "Springfield", "postal_code": "00000"}, "payment_methods": [ {"type": "card", "brand": "visa", "last4": "1111", "expires": "2029-08"} ] } ``` ### Clarifying Questions - Who consumes the output (analysts querying a warehouse, a product service reading one user's profile, or both), and how fresh must it be? - How much data arrives per batch or per second, and how quickly is it growing? - Does the batch source send full snapshots or only changed records, and does the stream carry full records or only changed fields? - Which fields are sensitive, who may see raw values, and must a user's deletion request reach every copy of their data? ### Part 1 — Batch upstream The upstream system delivers files of JSON records on a schedule. Describe the pipeline from file arrival to queryable output: where the raw data lands, how records are validated and transformed (including the nested address and payment fields), how results are loaded, and what happens when a run fails or the same file arrives twice. ```hint Make reruns boring Ask what must be true of each step so that rerunning yesterday's batch, or receiving the same file twice, leaves the output exactly as if it had run once. ``` #### What This Part Should Cover - Raw landing that keeps the input replayable, and detection of complete deliveries - Validation with quarantine of bad records, and flattening of the nested fields - Idempotent, partitioned loading with scheduling, retries and backfills ### Part 2 — Streaming upstream Now the upstream emits an event whenever a user is created or updated. How does the design change? ```hint What a stream takes away A stream gives you events for each user but never a moment when the data is complete. Think about duplicates, events that arrive out of order, and how a consumer resumes after a crash. ``` #### Clarifying Questions for this Part - Does each event carry an update timestamp or version number? - Does the upstream guarantee ordering for each user, or no ordering at all? #### What This Part Should Cover - Transport and partitioning that preserve the order of each user's events - Handling of duplicates, out-of-order and late events through idempotent upserts - Checkpointing, delivery guarantees and dead-lettering of bad events ### What a Strong Answer Covers - A brief, structured answer that matches the high-level scope of the question - A clear contrast between batch and streaming on latency, complexity, cost and reprocessing - Protection of names, addresses and payment data at every stage - Schema evolution of nested JSON and data-quality checks - Transformation logic shared by both paths instead of written twice ### Follow-up Questions - A user deletes their account and asks for their data to be erased. How does the deletion reach raw storage, processed tables and any streaming state? - The upstream adds a nested field or changes a field's type without notice. What happens in each design? - How would you backfill a year of history into the streaming design? - If a batch source and a streaming source must feed the same tables, how do you reconcile them?

Overview: Design a data processing system for JSON user records that contain names, addresses and payment methods, and explain how the design changes when the upstream source delivers data in batches versus as a continuous stream. It tests pipeline architecture, idempotent loading, ordering and late data, and the handling of sensitive fields.

Read the full Patreon Data Engineer interview experience this question came from

|Home/System Design/Patreon
Patreon logo
Patreon
Mar 14, 2026
easyData EngineerTechnical ScreenSystem Design
1
0

Your team receives user records as JSON. Each record holds personal data such as a name, a mailing address and one or more payment methods. The round opened with a short exercise in which the candidate processed a large JSON dump of such records using plain Python loops. The interviewer then asked a brief, high-level design question:

How would you build a data processing system for this data, and how would you build it if the upstream source delivers the data in batches, versus as a stream?

A short, structured answer was expected rather than a deep dive. The record below only illustrates the shape of the data; the exact fields were not reported.

{
  "id": "u-001",
  "name": "Sample User",
  "address": {"street": "1 Main St", "city": "Springfield", "postal_code": "00000"},
  "payment_methods": [
    {"type": "card", "brand": "visa", "last4": "1111", "expires": "2029-08"}
  ]
}

Clarifying Questions Guidance

  • Who consumes the output (analysts querying a warehouse, a product service reading one user's profile, or both), and how fresh must it be?
  • How much data arrives per batch or per second, and how quickly is it growing?
  • Does the batch source send full snapshots or only changed records, and does the stream carry full records or only changed fields?
  • Which fields are sensitive, who may see raw values, and must a user's deletion request reach every copy of their data?

Part 1 — Batch upstream

The upstream system delivers files of JSON records on a schedule. Describe the pipeline from file arrival to queryable output: where the raw data lands, how records are validated and transformed (including the nested address and payment fields), how results are loaded, and what happens when a run fails or the same file arrives twice.

What This Part Should Cover Guidance

  • Raw landing that keeps the input replayable, and detection of complete deliveries
  • Validation with quarantine of bad records, and flattening of the nested fields
  • Idempotent, partitioned loading with scheduling, retries and backfills

Part 2 — Streaming upstream

Now the upstream emits an event whenever a user is created or updated. How does the design change?

Clarifying Questions for this Part Guidance

  • Does each event carry an update timestamp or version number?
  • Does the upstream guarantee ordering for each user, or no ordering at all?

What This Part Should Cover Guidance

  • Transport and partitioning that preserve the order of each user's events
  • Handling of duplicates, out-of-order and late events through idempotent upserts
  • Checkpointing, delivery guarantees and dead-lettering of bad events

What a Strong Answer Covers Guidance

  • A brief, structured answer that matches the high-level scope of the question
  • A clear contrast between batch and streaming on latency, complexity, cost and reprocessing
  • Protection of names, addresses and payment data at every stage
  • Schema evolution of nested JSON and data-quality checks
  • Transformation logic shared by both paths instead of written twice

Follow-up Questions Guidance

  • A user deletes their account and asks for their data to be erased. How does the deletion reach raw storage, processed tables and any streaming state?
  • The upstream adds a nested field or changes a field's type without notice. What happens in each design?
  • How would you backfill a year of history into the streaming design?
  • If a batch source and a streaming source must feed the same tables, how do you reconcile them?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...