Design a Data Processing System for JSON User Records: Batch vs Streaming Input
Company: Patreon
Role: Data Engineer
Category: System Design
Difficulty: easy
Interview Round: Technical Screen
Your team receives user records as JSON. Each record holds personal data such as a name, a mailing address and one or more payment methods. The round opened with a short exercise in which the candidate processed a large JSON dump of such records using plain Python loops. The interviewer then asked a brief, high-level design question:
How would you build a data processing system for this data, and how would you build it if the upstream source delivers the data in batches, versus as a stream?
A short, structured answer was expected rather than a deep dive. The record below only illustrates the shape of the data; the exact fields were not reported.
```json
{
"id": "u-001",
"name": "Sample User",
"address": {"street": "1 Main St", "city": "Springfield", "postal_code": "00000"},
"payment_methods": [
{"type": "card", "brand": "visa", "last4": "1111", "expires": "2029-08"}
]
}
```
### Clarifying Questions
- Who consumes the output (analysts querying a warehouse, a product service reading one user's profile, or both), and how fresh must it be?
- How much data arrives per batch or per second, and how quickly is it growing?
- Does the batch source send full snapshots or only changed records, and does the stream carry full records or only changed fields?
- Which fields are sensitive, who may see raw values, and must a user's deletion request reach every copy of their data?
### Part 1 — Batch upstream
The upstream system delivers files of JSON records on a schedule. Describe the pipeline from file arrival to queryable output: where the raw data lands, how records are validated and transformed (including the nested address and payment fields), how results are loaded, and what happens when a run fails or the same file arrives twice.
```hint Make reruns boring
Ask what must be true of each step so that rerunning yesterday's batch, or receiving the same file twice, leaves the output exactly as if it had run once.
```
#### What This Part Should Cover
- Raw landing that keeps the input replayable, and detection of complete deliveries
- Validation with quarantine of bad records, and flattening of the nested fields
- Idempotent, partitioned loading with scheduling, retries and backfills
### Part 2 — Streaming upstream
Now the upstream emits an event whenever a user is created or updated. How does the design change?
```hint What a stream takes away
A stream gives you events for each user but never a moment when the data is complete. Think about duplicates, events that arrive out of order, and how a consumer resumes after a crash.
```
#### Clarifying Questions for this Part
- Does each event carry an update timestamp or version number?
- Does the upstream guarantee ordering for each user, or no ordering at all?
#### What This Part Should Cover
- Transport and partitioning that preserve the order of each user's events
- Handling of duplicates, out-of-order and late events through idempotent upserts
- Checkpointing, delivery guarantees and dead-lettering of bad events
### What a Strong Answer Covers
- A brief, structured answer that matches the high-level scope of the question
- A clear contrast between batch and streaming on latency, complexity, cost and reprocessing
- Protection of names, addresses and payment data at every stage
- Schema evolution of nested JSON and data-quality checks
- Transformation logic shared by both paths instead of written twice
### Follow-up Questions
- A user deletes their account and asks for their data to be erased. How does the deletion reach raw storage, processed tables and any streaming state?
- The upstream adds a nested field or changes a field's type without notice. What happens in each design?
- How would you backfill a year of history into the streaming design?
- If a batch source and a streaming source must feed the same tables, how do you reconcile them?
Overview: Design a data processing system for JSON user records that contain names, addresses and payment methods, and explain how the design changes when the upstream source delivers data in batches versus as a continuous stream. It tests pipeline architecture, idempotent loading, ordering and late data, and the handling of sensitive fields.
Read the full Patreon Data Engineer interview experience this question came from