Design a Product Catalog Ingested From Full Retailer Snapshots With Admin Overrides
Company: Attentive
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a product catalog system. Retailers send their catalogs to the platform as snapshot files, users browse the resulting catalog, and platform admins can override product attributes.
After clarifying with the interviewer, the requirements are:
**Ingestion**
- Retailers provide catalog data as snapshot files. Each file contains the retailer's full catalog, not an incremental change set.
- There are 50,000 retailers. The average retailer has 20,000 products, but the median retailer has 50.
- No versioning is required. While an update is being applied product by product, a user query may see some products already updated and others not yet updated, and that is acceptable.
**Browsing**
- Users can list all product SKUs of a retailer by `retailer_id`.
- Users can get the details of a product by `product_id`.
- There are 50 million users and 11,000 queries per second, and queries must be low latency, within 400 ms.
**Admin overrides**
- Admins can override any attribute of any product, system-wide.
- A retailer's update must not change an overridden value.
- A retailer can delete a product even if an admin has overridden some of its attribute values.
- An override can be deleted. Once it is deleted, users must see the retailer's newest value for that attribute.
Start with the core design: the entities, the database choice, and how data flows in and out. The interviewer then asks four follow-up questions, given below as Parts 2 to 5.
### Clarifying Questions
- Is `product_id` globally unique, or unique only within a retailer, and how does it relate to the retailer's SKU?
- How often do existing retailers upload snapshots, and how large is the largest retailer's file? Only the average and the median product counts are given.
- If a retailer deletes a product that has overrides and later sends the same product again, should the old overrides apply again?
- Is the 400 ms target a percentile such as p99, and does the SKU listing need pagination for the largest retailers?
### Part 1 — Core design
Propose the entities, the database choice, and the ingestion and read paths that meet the requirements above, including all of the override rules.
```hint Three events, one model
Write down what must happen to an overridden attribute when the retailer updates the product, when the retailer deletes it, and when the admin deletes the override. Look for a data model in which all three outcomes follow without special cases.
```
#### What This Part Should Cover
- A data model in which retailer values and admin overrides cannot clobber each other, and every override rule holds
- Storage and caching choices justified against 11,000 QPS, the 400 ms target, and the size skew across retailers
- The flow from an uploaded file to visible products, and the two read APIs
### Part 2 — Comparing two snapshot files
How do you compare two snapshot files to find what changed?
```hint Compare against what
A snapshot is a set of rows keyed by SKU. Decide what you compare the new file with, and what you need per row to classify it without comparing every attribute.
```
#### What This Part Should Cover
- How added, removed, changed and unchanged products are detected, and the time and memory cost for the largest retailers
- What the new snapshot is compared against, and why that choice matters when an earlier apply did not finish
- Normalization that prevents formatting noise from counting as a change
### Part 3 — A wrong snapshot, quickly replaced
A retailer uploads snapshot 1, quickly realizes it is wrong, and uploads snapshot 2. The catalog must end up showing the complete snapshot 2. How do you make sure of that?
```hint Two writers, one retailer
Picture both snapshots being processed at the same time. Ask what each write would need to know to avoid overwriting newer data, and whether the result for snapshot 2 should depend on snapshot 1 having finished.
```
#### Clarifying Questions for this Part
- Can snapshot 2 arrive while snapshot 1 is still being applied, or even before processing of snapshot 1 has started?
#### What This Part Should Cover
- An ordering of snapshots per retailer that does not depend on processing timing
- Behavior when snapshot 2 arrives before, during, and after the processing of snapshot 1
- Why the final state equals snapshot 2 exactly, including products that snapshot 1 wrongly removed or changed, and their overrides
### Part 4 — 50,000 new retailers uploading daily
Now 50,000 new retailers join, and each of them uploads a snapshot every day. How does the system handle this?
```hint Mind the skew
The median retailer has 50 products while the average has 20,000. Ask what that skew does to a single shared ingestion queue, and what a daily upload that changes almost nothing should cost.
```
#### What This Part Should Cover
- A load estimate for ingestion that separates rows read and compared from rows actually written
- Horizontal scaling of uploads, ingestion workers, and storage
- Fairness and isolation between very large and very small retailers
### Part 5 — Observability
How would you build observability for this system? What measures would you take to make sure everything runs as expected?
```hint Healthy is not correct
Separate signals that the system is up from signals that the catalog is right. Which checks would prove that a snapshot was applied fully and faithfully?
```
#### What This Part Should Cover
- Metrics and service-level objectives for both the read path and the ingestion pipeline
- Data-correctness checks, not only health checks
- Alerting, tracing, and an audit trail for overrides
### What a Strong Answer Covers
- Requirements restated with numbers, including what the gap between the average and the median implies
- Overrides modeled so that every stated rule follows from the data model rather than from special cases
- Ingestion of full snapshots that is ordered per retailer, idempotent, and self-correcting
- A read path whose latency stays comfortably within 400 ms at 11,000 QPS, isolated from ingestion load
- Concrete failure modes, and how each one is detected
### Follow-up Questions
- If users must never see a half-applied snapshot, what would you change, and what would it cost?
- How would you handle a single retailer whose snapshot takes hours to diff and apply?
- How would admins find overrides whose underlying retailer values have since changed, so stale overrides can be reviewed?
Overview: A system design question about a product catalog built from full retailer snapshot files, with 50,000 retailers, 11,000 read queries per second and a 400 ms latency target. It tests snapshot diffing, recovery from a mistaken upload, admin overrides that survive retailer updates, scaling ingestion, and observability.
Read the full Attentive Software Engineer interview experience this question came from