Design a Large-Scale Personalized Recommendation System End to End

Read the full interview experience this question came from →

Quick Overview

An ML system design question asking you to design a large-scale personalized recommendation system end to end, from candidate retrieval and multi-stage ranking to training data, evaluation, serving and monitoring. It tests problem framing, latency budgeting, cold start, feedback loops and online experimentation.

Design a Large-Scale Personalized Recommendation System End to End

Company: Meta

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

Design a recommendation system end to end. When a user opens a product surface, the system picks a small, ordered set of items to show out of a very large and constantly changing catalog, and it keeps learning from how users respond. Cover the whole loop: how the problem is framed as a machine learning problem, how candidates are retrieved and ranked, what data and features the models use, how they are trained and evaluated, how the system serves requests within its latency budget, and how it is monitored. ```hint The catalog is too big to score The most accurate model cannot score every item in the catalog for every request in time. Decide how to divide the work between cheap and expensive stages, and how many items each stage passes on. ``` ```hint Your data comes from your own system Users can only react to items the system chose to show them. Think about what that does to your training labels, to offline evaluation, and to new items. ``` ### Constraints and Clarifications - This was the system design round of an onsite for a machine learning engineer working on search, ads and recommendations. - The item type, product surface, scale and success metric are not given. Ask for them or state your assumptions. ### Clarifying Questions - What is being recommended (posts, videos, products, accounts to follow), and on which surface (a home feed, related items, notifications)? - What is the business goal (time spent, meaningful interactions, purchases, long-term retention), and which guardrails matter, such as integrity, diversity or fairness to creators? - How many users, items and requests per second, and what is the latency budget per request? - How quickly must new items become recommendable: within minutes, or within a day? - Is there logged interaction history to train on? ### What a Strong Answer Covers - Framing: a business objective turned into prediction targets, labels and a final ranking score - A multi-stage architecture with the candidate counts and latency budget of each stage - Candidate sources and ranking models, with the features and training data each one uses - Offline metrics and online experiments, including the biases in logged feedback - Serving at scale, freshness, cold start for new users and new items, and feedback loops - Monitoring and fallbacks for data, models and serving ### Follow-up Questions - A new item has no interactions yet. How does it get its first impressions, and how do you keep it from being buried? - Offline metrics improve but the online A/B test shows no gain. What are the likely causes, and how do you investigate? - How would you combine several objectives, such as clicks, watch time, shares and hides, into one ranking? - How do you keep the feed from narrowing each user's interests over time?

Overview: An ML system design question asking you to design a large-scale personalized recommendation system end to end, from candidate retrieval and multi-stage ranking to training data, evaluation, serving and monitoring. It tests problem framing, latency budgeting, cold start, feedback loops and online experimentation.

Read the full Meta Machine Learning Engineer interview experience this question came from

|Home/ML System Design/Meta
Meta logo
Meta
Jul 7, 2026
mediumMachine Learning EngineerOnsiteML System Design
0
0

Design a recommendation system end to end. When a user opens a product surface, the system picks a small, ordered set of items to show out of a very large and constantly changing catalog, and it keeps learning from how users respond.

Cover the whole loop: how the problem is framed as a machine learning problem, how candidates are retrieved and ranked, what data and features the models use, how they are trained and evaluated, how the system serves requests within its latency budget, and how it is monitored.

Constraints and Clarifications

  • This was the system design round of an onsite for a machine learning engineer working on search, ads and recommendations.
  • The item type, product surface, scale and success metric are not given. Ask for them or state your assumptions.

Clarifying Questions Guidance

  • What is being recommended (posts, videos, products, accounts to follow), and on which surface (a home feed, related items, notifications)?
  • What is the business goal (time spent, meaningful interactions, purchases, long-term retention), and which guardrails matter, such as integrity, diversity or fairness to creators?
  • How many users, items and requests per second, and what is the latency budget per request?
  • How quickly must new items become recommendable: within minutes, or within a day?
  • Is there logged interaction history to train on?

What a Strong Answer Covers Guidance

  • Framing: a business objective turned into prediction targets, labels and a final ranking score
  • A multi-stage architecture with the candidate counts and latency budget of each stage
  • Candidate sources and ranking models, with the features and training data each one uses
  • Offline metrics and online experiments, including the biases in logged feedback
  • Serving at scale, freshness, cold start for new users and new items, and feedback loops
  • Monitoring and fallbacks for data, models and serving

Follow-up Questions Guidance

  • A new item has no interactions yet. How does it get its first impressions, and how do you keep it from being buried?
  • Offline metrics improve but the online A/B test shows no gain. What are the likely causes, and how do you investigate?
  • How would you combine several objectives, such as clicks, watch time, shares and hides, into one ranking?
  • How do you keep the feed from narrowing each user's interests over time?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...