Design a watch-next video recommender and derive its losses and optimizer update rules
Company: Google
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: medium
Interview Round: Onsite
Design the "watch next" recommendations for a large video-sharing platform. While a user is watching a video, the system chooses and orders a list of videos to offer as what to watch next.
This interview goes deep on modeling details. Beyond the architecture, expect to write down the exact loss function of every model you propose and the update formulas of the optimizer you would train it with.
### Clarifying Questions
- Is the surface a ranked list shown next to the video, a single video that plays automatically when the current one ends, or both?
- What is the primary objective: watch time, user satisfaction, or a blend, and which guardrail metrics must not regress?
- Is the user signed in, so that watch history is available?
- What latency budget applies to producing the list?
- How large is the catalog, and how quickly do new videos need to become recommendable?
### Part 1 — Design the watch-next system
Frame the problem, then design the system end to end: the objective and training labels, how candidates are generated, how they are ranked, how the list is served, and how you evaluate it offline and online.
```hint What counts as a good next video
A click on a recommendation is easy to measure and easy to game. Decide which signal of a good next video you actually want to optimize.
```
```hint The video on screen
The video being watched right now is a strong signal you have on every request. Decide where it enters the system.
```
#### What This Part Should Cover
- An objective beyond clicks, and labels that match it
- A multi-stage design (candidate generation, ranking, re-ranking) where the current video drives both retrieval and ranking
- Features, serving path and latency
- Offline metrics, online experiments, and monitoring for feedback loops
### Part 2 — Write the loss functions
For each model in your design, write the loss you would train it with, define every symbol, and explain why it fits the label. Include how you would combine several objectives into one training loss.
```hint One loss per kind of label
Binary outcomes, continuous outcomes such as watch time, and picking one item out of a huge catalog each call for a different loss.
```
```hint Negatives nobody logged
A retrieval model sees mostly positives in the logs. Decide where its negatives come from, and what that choice does to the loss.
```
#### What This Part Should Cover
- Correct formulas with every symbol defined
- A retrieval loss over sampled negatives, and the bias that sampling introduces
- Ranking losses for engagement and for watch time
- Combining objectives, and handling position bias in the training data
### Part 3 — Write the optimizer updates
Write the update rule of the optimizer you would use, starting from plain stochastic gradient descent and explaining each added term. Show the gradient of your ranking loss with respect to the model's output logit.
```hint Build it up
Write the plain gradient step first, then add one idea at a time: memory of past gradients, then per-parameter step sizes.
```
```hint Sparse and dense parameters
Embedding rows for rarely seen videos receive gradients only occasionally. Consider whether one step size for every parameter suits them.
```
#### What This Part Should Cover
- Correct update rules, including bias correction where it applies
- A correct gradient derivation for the ranking loss
- A justified optimizer choice for embedding tables versus dense layers
- Practical settings such as learning-rate schedules, gradient clipping and weight decay
### What a Strong Answer Covers
- A clear objective beyond clicks, with labels that match it
- A multi-stage system in which the current video drives candidate generation and ranking
- Exact, correctly defined losses for each label, including sampled negatives and position bias
- Correct optimizer update rules, with justified choices for sparse and dense parameters
- Offline metrics that predict online results, an A/B test plan, and monitoring for feedback loops
### Follow-up Questions
- Your retrieval model keeps recommending the same popular videos. Which part of your loss is responsible, and how do you fix it?
- How would you make a video uploaded an hour ago eligible for watch-next?
- Why does Adam with L2 regularization in the loss behave differently from Adam with decoupled weight decay, and when does it matter?
- How do you keep the system from maximizing watch time with ever more sensational videos?
Overview: Design the watch-next recommendations shown while a user watches a video, from objectives and candidate generation to ranking, serving and evaluation. The interview probes modeling depth: exact loss functions for retrieval and multi-task ranking, the gradient at the output logit, and the update rules of the optimizers used to train each model.