Review a Feature Pipeline PR for Production Risks, Then Serve Features by Freshness
Company: Netflix
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Technical Screen
In the practical coding half of a machine learning screen, the interviewer pastes a pull request that changes a feature pipeline: the code that turns raw event data into features for a model. Every test in the PR passes. You are asked to review it live and explain what problems the change could cause if it lands in production. The conversation then moves on to how these features should be served when the model needs them at different levels of freshness.
### Constraints and Clarifications
- The PR's code is not reproduced here. Practice by naming what you would check in such a PR, what each problem would look like in the code, and how it would surface in production. If you practice with a partner, have them write a short feature-computation function for you to review.
- "Tests pass" means the logic is correct on the test fixtures. The question is about everything those fixtures do not exercise.
### Clarifying Questions
- Are these features used for training, for online inference, or both? If both, is the same code used to compute them offline and online?
- What are the inputs (event streams, warehouse tables), how large are they, and how late can events arrive?
- Is the pipeline a scheduled batch job, a streaming job, or code that runs at request time?
- What happens downstream when a feature value is missing or stale?
### Part 1 — Production risks in a PR whose tests pass
Review the PR. What could go wrong once it runs in production, even though its tests pass? Rank the problems by impact, and for each one, say how you would catch it before release or detect it afterwards.
```hint What the fixtures lack
Test fixtures are small, clean and frozen in time. List the properties of real production data that such fixtures usually lack.
```
```hint Time is a column too
For every aggregate in the PR, ask which events it is allowed to see when the model makes a prediction at a given moment.
```
```hint Same name, same meaning
Ask whether the value the model sees at serving time is computed the same way, from the same data, as the value it was trained on.
```
#### What This Part Should Cover
- Correctness at each point in time: leakage of future data, late and out-of-order events, timezone and window boundaries
- Data quality: nulls, duplicates, join fan-out, and schema or unit changes upstream
- Consistency between training and serving, and the rollout of a changed feature definition
- Operational risks (skewed keys, memory, runtime, idempotent reruns and backfills) and the checks and monitoring that catch them
### Part 2 — Serving the features at different freshness requirements
How would you serve these features if the model can use values that are a day old, if they must reflect events from the last few minutes, or if they depend on the request itself? Explain what changes in the architecture, the consistency guarantees, the latency and the cost.
```hint Before or during the request
For each freshness level, decide whether the value is computed before the request arrives or while it is being served, and what sits between the computation and the model.
```
```hint One definition, two paths
If a feature is computed by a streaming job for serving and by a batch job for training, think about what keeps the two from drifting apart.
```
#### What This Part Should Cover
- Precomputed batch features in an online store, streaming aggregates, and on-demand features computed at request time
- Latency, cost and operational complexity of each option
- Keeping the values used in training consistent with the values served online
- Staleness, missing values and fallbacks when the fresh path fails
### What a Strong Answer Covers
- Concrete, prioritized risks specific to feature pipelines rather than code-style comments
- Leakage and training-serving skew identified as the most damaging failures, because they are silent
- A serving design matched to how much freshness each feature actually needs
- A rollout and monitoring plan that would catch silent data problems
### Follow-up Questions
- The PR changes the definition of a feature that the deployed model already uses. How do you roll it out safely?
- How would you detect training-serving skew in production?
- The streaming job that maintains a recent-activity feature falls an hour behind. What should the model see, and how should the system behave?
- Which tests would you add to the PR so that the problems you found would have failed CI?
Overview: Review a feature pipeline pull request whose tests all pass and explain what could break once it runs in production, from time leakage and training-serving skew to data quality and scale. Then decide how the same features should be served when the model needs day-old, minute-fresh or request-time values.
Read the full Netflix Machine Learning Engineer interview experience this question came from