Design a Product Catalog Ingested From Full Retailer Snapshots With Admin Overrides

Read the full interview experience this question came from →

Quick Overview

A system design question about a product catalog built from full retailer snapshot files, with 50,000 retailers, 11,000 read queries per second and a 400 ms latency target. It tests snapshot diffing, recovery from a mistaken upload, admin overrides that survive retailer updates, scaling ingestion, and observability.

Design a Product Catalog Ingested From Full Retailer Snapshots With Admin Overrides

Company: Attentive

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a product catalog system. Retailers send their catalogs to the platform as snapshot files, users browse the resulting catalog, and platform admins can override product attributes. After clarifying with the interviewer, the requirements are: **Ingestion** - Retailers provide catalog data as snapshot files. Each file contains the retailer's full catalog, not an incremental change set. - There are 50,000 retailers. The average retailer has 20,000 products, but the median retailer has 50. - No versioning is required. While an update is being applied product by product, a user query may see some products already updated and others not yet updated, and that is acceptable. **Browsing** - Users can list all product SKUs of a retailer by `retailer_id`. - Users can get the details of a product by `product_id`. - There are 50 million users and 11,000 queries per second, and queries must be low latency, within 400 ms. **Admin overrides** - Admins can override any attribute of any product, system-wide. - A retailer's update must not change an overridden value. - A retailer can delete a product even if an admin has overridden some of its attribute values. - An override can be deleted. Once it is deleted, users must see the retailer's newest value for that attribute. Start with the core design: the entities, the database choice, and how data flows in and out. The interviewer then asks four follow-up questions, given below as Parts 2 to 5. ### Clarifying Questions - Is `product_id` globally unique, or unique only within a retailer, and how does it relate to the retailer's SKU? - How often do existing retailers upload snapshots, and how large is the largest retailer's file? Only the average and the median product counts are given. - If a retailer deletes a product that has overrides and later sends the same product again, should the old overrides apply again? - Is the 400 ms target a percentile such as p99, and does the SKU listing need pagination for the largest retailers? ### Part 1 — Core design Propose the entities, the database choice, and the ingestion and read paths that meet the requirements above, including all of the override rules. ```hint Three events, one model Write down what must happen to an overridden attribute when the retailer updates the product, when the retailer deletes it, and when the admin deletes the override. Look for a data model in which all three outcomes follow without special cases. ``` #### What This Part Should Cover - A data model in which retailer values and admin overrides cannot clobber each other, and every override rule holds - Storage and caching choices justified against 11,000 QPS, the 400 ms target, and the size skew across retailers - The flow from an uploaded file to visible products, and the two read APIs ### Part 2 — Comparing two snapshot files How do you compare two snapshot files to find what changed? ```hint Compare against what A snapshot is a set of rows keyed by SKU. Decide what you compare the new file with, and what you need per row to classify it without comparing every attribute. ``` #### What This Part Should Cover - How added, removed, changed and unchanged products are detected, and the time and memory cost for the largest retailers - What the new snapshot is compared against, and why that choice matters when an earlier apply did not finish - Normalization that prevents formatting noise from counting as a change ### Part 3 — A wrong snapshot, quickly replaced A retailer uploads snapshot 1, quickly realizes it is wrong, and uploads snapshot 2. The catalog must end up showing the complete snapshot 2. How do you make sure of that? ```hint Two writers, one retailer Picture both snapshots being processed at the same time. Ask what each write would need to know to avoid overwriting newer data, and whether the result for snapshot 2 should depend on snapshot 1 having finished. ``` #### Clarifying Questions for this Part - Can snapshot 2 arrive while snapshot 1 is still being applied, or even before processing of snapshot 1 has started? #### What This Part Should Cover - An ordering of snapshots per retailer that does not depend on processing timing - Behavior when snapshot 2 arrives before, during, and after the processing of snapshot 1 - Why the final state equals snapshot 2 exactly, including products that snapshot 1 wrongly removed or changed, and their overrides ### Part 4 — 50,000 new retailers uploading daily Now 50,000 new retailers join, and each of them uploads a snapshot every day. How does the system handle this? ```hint Mind the skew The median retailer has 50 products while the average has 20,000. Ask what that skew does to a single shared ingestion queue, and what a daily upload that changes almost nothing should cost. ``` #### What This Part Should Cover - A load estimate for ingestion that separates rows read and compared from rows actually written - Horizontal scaling of uploads, ingestion workers, and storage - Fairness and isolation between very large and very small retailers ### Part 5 — Observability How would you build observability for this system? What measures would you take to make sure everything runs as expected? ```hint Healthy is not correct Separate signals that the system is up from signals that the catalog is right. Which checks would prove that a snapshot was applied fully and faithfully? ``` #### What This Part Should Cover - Metrics and service-level objectives for both the read path and the ingestion pipeline - Data-correctness checks, not only health checks - Alerting, tracing, and an audit trail for overrides ### What a Strong Answer Covers - Requirements restated with numbers, including what the gap between the average and the median implies - Overrides modeled so that every stated rule follows from the data model rather than from special cases - Ingestion of full snapshots that is ordered per retailer, idempotent, and self-correcting - A read path whose latency stays comfortably within 400 ms at 11,000 QPS, isolated from ingestion load - Concrete failure modes, and how each one is detected ### Follow-up Questions - If users must never see a half-applied snapshot, what would you change, and what would it cost? - How would you handle a single retailer whose snapshot takes hours to diff and apply? - How would admins find overrides whose underlying retailer values have since changed, so stale overrides can be reviewed?

Overview: A system design question about a product catalog built from full retailer snapshot files, with 50,000 retailers, 11,000 read queries per second and a 400 ms latency target. It tests snapshot diffing, recovery from a mistaken upload, admin overrides that survive retailer updates, scaling ingestion, and observability.

Read the full Attentive Software Engineer interview experience this question came from

|Home/System Design/Attentive
Attentive logo
Attentive
Sep 24, 2026
mediumSoftware EngineerOnsiteSystem Design
0
0

Design a product catalog system. Retailers send their catalogs to the platform as snapshot files, users browse the resulting catalog, and platform admins can override product attributes.

After clarifying with the interviewer, the requirements are:

Ingestion

  • Retailers provide catalog data as snapshot files. Each file contains the retailer's full catalog, not an incremental change set.
  • There are 50,000 retailers. The average retailer has 20,000 products, but the median retailer has 50.
  • No versioning is required. While an update is being applied product by product, a user query may see some products already updated and others not yet updated, and that is acceptable.

Browsing

  • Users can list all product SKUs of a retailer by retailer_id .
  • Users can get the details of a product by product_id .
  • There are 50 million users and 11,000 queries per second, and queries must be low latency, within 400 ms.

Admin overrides

  • Admins can override any attribute of any product, system-wide.
  • A retailer's update must not change an overridden value.
  • A retailer can delete a product even if an admin has overridden some of its attribute values.
  • An override can be deleted. Once it is deleted, users must see the retailer's newest value for that attribute.

Start with the core design: the entities, the database choice, and how data flows in and out. The interviewer then asks four follow-up questions, given below as Parts 2 to 5.

Clarifying Questions Guidance

  • Is product_id globally unique, or unique only within a retailer, and how does it relate to the retailer's SKU?
  • How often do existing retailers upload snapshots, and how large is the largest retailer's file? Only the average and the median product counts are given.
  • If a retailer deletes a product that has overrides and later sends the same product again, should the old overrides apply again?
  • Is the 400 ms target a percentile such as p99, and does the SKU listing need pagination for the largest retailers?

Part 1 — Core design

Propose the entities, the database choice, and the ingestion and read paths that meet the requirements above, including all of the override rules.

What This Part Should Cover Guidance

  • A data model in which retailer values and admin overrides cannot clobber each other, and every override rule holds
  • Storage and caching choices justified against 11,000 QPS, the 400 ms target, and the size skew across retailers
  • The flow from an uploaded file to visible products, and the two read APIs

Part 2 — Comparing two snapshot files

How do you compare two snapshot files to find what changed?

What This Part Should Cover Guidance

  • How added, removed, changed and unchanged products are detected, and the time and memory cost for the largest retailers
  • What the new snapshot is compared against, and why that choice matters when an earlier apply did not finish
  • Normalization that prevents formatting noise from counting as a change

Part 3 — A wrong snapshot, quickly replaced

A retailer uploads snapshot 1, quickly realizes it is wrong, and uploads snapshot 2. The catalog must end up showing the complete snapshot 2. How do you make sure of that?

Clarifying Questions for this Part Guidance

  • Can snapshot 2 arrive while snapshot 1 is still being applied, or even before processing of snapshot 1 has started?

What This Part Should Cover Guidance

  • An ordering of snapshots per retailer that does not depend on processing timing
  • Behavior when snapshot 2 arrives before, during, and after the processing of snapshot 1
  • Why the final state equals snapshot 2 exactly, including products that snapshot 1 wrongly removed or changed, and their overrides

Part 4 — 50,000 new retailers uploading daily

Now 50,000 new retailers join, and each of them uploads a snapshot every day. How does the system handle this?

What This Part Should Cover Guidance

  • A load estimate for ingestion that separates rows read and compared from rows actually written
  • Horizontal scaling of uploads, ingestion workers, and storage
  • Fairness and isolation between very large and very small retailers

Part 5 — Observability

How would you build observability for this system? What measures would you take to make sure everything runs as expected?

What This Part Should Cover Guidance

  • Metrics and service-level objectives for both the read path and the ingestion pipeline
  • Data-correctness checks, not only health checks
  • Alerting, tracing, and an audit trail for overrides

What a Strong Answer Covers Guidance

  • Requirements restated with numbers, including what the gap between the average and the median implies
  • Overrides modeled so that every stated rule follows from the data model rather than from special cases
  • Ingestion of full snapshots that is ordered per retailer, idempotent, and self-correcting
  • A read path whose latency stays comfortably within 400 ms at 11,000 QPS, isolated from ingestion load
  • Concrete failure modes, and how each one is detected

Follow-up Questions Guidance

  • If users must never see a half-applied snapshot, what would you change, and what would it cost?
  • How would you handle a single retailer whose snapshot takes hours to diff and apply?
  • How would admins find overrides whose underlying retailer values have since changed, so stale overrides can be reviewed?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...