Design fast query-to-listing retrieval over tens of millions of listings

Read the full interview experience this question came from →

Quick Overview

An ML system design question on retrieving relevant listings for free-text search queries across a catalog of tens of millions of items with low latency. It tests candidate generation, indexing, embedding training, freshness, filtering and scaling, plus the offline and A/B test metrics used to evaluate retrieval.

Design fast query-to-listing retrieval over tens of millions of listings

Company: eBay

Role: Applied Scientist

Category: ML System Design

Difficulty: easy

Interview Round: Technical Screen

Users type free-text search queries into a marketplace, and the catalog holds tens of millions of listings. Design how the system quickly retrieves the listings relevant to a query. Expect the discussion to drill into each component you propose. ```hint Precompute Think about what can be computed for every listing before any query arrives. ``` ```hint Complementary signals Exact term matching and semantic matching each fail on different kinds of queries. ``` ### Constraints and Clarifications - The catalog holds tens of millions of listings. No latency budget, query volume, or freshness requirement was specified; ask for them. - Assume listings have at least a title, description, category, structured attributes, price and availability. ### Clarifying Questions - What latency budget does retrieval get within the full search request, and what peak query rate must it sustain? - How quickly must a new, edited, sold or removed listing be reflected in results? - Does retrieval produce a candidate set for a downstream ranker, or the final ordered results? - Which hard filters must be honored (category, price range, location, condition, availability)? - What interaction logs exist (queries, impressions, clicks, purchases) for training and evaluation? ### What a Strong Answer Covers - Requirements and a clear definition of relevance for this marketplace - A retrieval architecture that avoids scoring every listing per query, combining lexical and semantic candidate generation before ranking - Offline pipelines: embedding model training from interaction data, index building, and incremental updates - Serving: sharding, replication, filter handling, latency budgeting and graceful degradation - Evaluation with offline retrieval metrics and online A/B test metrics, plus monitoring ### Follow-up Questions - How do you train the query and listing encoders, and where do the negative examples come from? - A listing sells out. How quickly, and through what path, does it disappear from both the lexical and the vector index? - How do you apply a price or category filter with approximate nearest-neighbor search without losing recall? - Which offline metrics and which A/B test metrics would decide whether a new retrieval model ships?

Overview: An ML system design question on retrieving relevant listings for free-text search queries across a catalog of tens of millions of items with low latency. It tests candidate generation, indexing, embedding training, freshness, filtering and scaling, plus the offline and A/B test metrics used to evaluate retrieval.

Read the full eBay Applied Scientist interview experience this question came from

|Home/ML System Design/eBay
eBay logo
eBay
Sep 24, 2026
easyApplied ScientistTechnical ScreenML System Design
0
0

Users type free-text search queries into a marketplace, and the catalog holds tens of millions of listings. Design how the system quickly retrieves the listings relevant to a query. Expect the discussion to drill into each component you propose.

Constraints and Clarifications

  • The catalog holds tens of millions of listings. No latency budget, query volume, or freshness requirement was specified; ask for them.
  • Assume listings have at least a title, description, category, structured attributes, price and availability.

Clarifying Questions Guidance

  • What latency budget does retrieval get within the full search request, and what peak query rate must it sustain?
  • How quickly must a new, edited, sold or removed listing be reflected in results?
  • Does retrieval produce a candidate set for a downstream ranker, or the final ordered results?
  • Which hard filters must be honored (category, price range, location, condition, availability)?
  • What interaction logs exist (queries, impressions, clicks, purchases) for training and evaluation?

What a Strong Answer Covers Guidance

  • Requirements and a clear definition of relevance for this marketplace
  • A retrieval architecture that avoids scoring every listing per query, combining lexical and semantic candidate generation before ranking
  • Offline pipelines: embedding model training from interaction data, index building, and incremental updates
  • Serving: sharding, replication, filter handling, latency budgeting and graceful degradation
  • Evaluation with offline retrieval metrics and online A/B test metrics, plus monitoring

Follow-up Questions Guidance

  • How do you train the query and listing encoders, and where do the negative examples come from?
  • A listing sells out. How quickly, and through what path, does it disappear from both the lexical and the vector index?
  • How do you apply a price or category filter with approximate nearest-neighbor search without losing recall?
  • Which offline metrics and which A/B test metrics would decide whether a new retrieval model ships?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...