Design fast query-to-listing retrieval over tens of millions of listings
Company: eBay
Role: Applied Scientist
Category: ML System Design
Difficulty: easy
Interview Round: Technical Screen
Users type free-text search queries into a marketplace, and the catalog holds tens of millions of listings. Design how the system quickly retrieves the listings relevant to a query. Expect the discussion to drill into each component you propose.
```hint Precompute
Think about what can be computed for every listing before any query arrives.
```
```hint Complementary signals
Exact term matching and semantic matching each fail on different kinds of queries.
```
### Constraints and Clarifications
- The catalog holds tens of millions of listings. No latency budget, query volume, or freshness requirement was specified; ask for them.
- Assume listings have at least a title, description, category, structured attributes, price and availability.
### Clarifying Questions
- What latency budget does retrieval get within the full search request, and what peak query rate must it sustain?
- How quickly must a new, edited, sold or removed listing be reflected in results?
- Does retrieval produce a candidate set for a downstream ranker, or the final ordered results?
- Which hard filters must be honored (category, price range, location, condition, availability)?
- What interaction logs exist (queries, impressions, clicks, purchases) for training and evaluation?
### What a Strong Answer Covers
- Requirements and a clear definition of relevance for this marketplace
- A retrieval architecture that avoids scoring every listing per query, combining lexical and semantic candidate generation before ranking
- Offline pipelines: embedding model training from interaction data, index building, and incremental updates
- Serving: sharding, replication, filter handling, latency budgeting and graceful degradation
- Evaluation with offline retrieval metrics and online A/B test metrics, plus monitoring
### Follow-up Questions
- How do you train the query and listing encoders, and where do the negative examples come from?
- A listing sells out. How quickly, and through what path, does it disappear from both the lexical and the vector index?
- How do you apply a price or category filter with approximate nearest-neighbor search without losing recall?
- Which offline metrics and which A/B test metrics would decide whether a new retrieval model ships?
Overview: An ML system design question on retrieving relevant listings for free-text search queries across a catalog of tens of millions of items with low latency. It tests candidate generation, indexing, embedding training, freshness, filtering and scaling, plus the offline and A/B test metrics used to evaluate retrieval.
Read the full eBay Applied Scientist interview experience this question came from