Review a Feature Pipeline PR for Production Risks, Then Serve Features by Freshness

Read the full interview experience this question came from →

Quick Overview

Review a feature pipeline pull request whose tests all pass and explain what could break once it runs in production, from time leakage and training-serving skew to data quality and scale. Then decide how the same features should be served when the model needs day-old, minute-fresh or request-time values.

Review a Feature Pipeline PR for Production Risks, Then Serve Features by Freshness

Company: Netflix

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: hard

Interview Round: Technical Screen

In the practical coding half of a machine learning screen, the interviewer pastes a pull request that changes a feature pipeline: the code that turns raw event data into features for a model. Every test in the PR passes. You are asked to review it live and explain what problems the change could cause if it lands in production. The conversation then moves on to how these features should be served when the model needs them at different levels of freshness. ### Constraints and Clarifications - The PR's code is not reproduced here. Practice by naming what you would check in such a PR, what each problem would look like in the code, and how it would surface in production. If you practice with a partner, have them write a short feature-computation function for you to review. - "Tests pass" means the logic is correct on the test fixtures. The question is about everything those fixtures do not exercise. ### Clarifying Questions - Are these features used for training, for online inference, or both? If both, is the same code used to compute them offline and online? - What are the inputs (event streams, warehouse tables), how large are they, and how late can events arrive? - Is the pipeline a scheduled batch job, a streaming job, or code that runs at request time? - What happens downstream when a feature value is missing or stale? ### Part 1 — Production risks in a PR whose tests pass Review the PR. What could go wrong once it runs in production, even though its tests pass? Rank the problems by impact, and for each one, say how you would catch it before release or detect it afterwards. ```hint What the fixtures lack Test fixtures are small, clean and frozen in time. List the properties of real production data that such fixtures usually lack. ``` ```hint Time is a column too For every aggregate in the PR, ask which events it is allowed to see when the model makes a prediction at a given moment. ``` ```hint Same name, same meaning Ask whether the value the model sees at serving time is computed the same way, from the same data, as the value it was trained on. ``` #### What This Part Should Cover - Correctness at each point in time: leakage of future data, late and out-of-order events, timezone and window boundaries - Data quality: nulls, duplicates, join fan-out, and schema or unit changes upstream - Consistency between training and serving, and the rollout of a changed feature definition - Operational risks (skewed keys, memory, runtime, idempotent reruns and backfills) and the checks and monitoring that catch them ### Part 2 — Serving the features at different freshness requirements How would you serve these features if the model can use values that are a day old, if they must reflect events from the last few minutes, or if they depend on the request itself? Explain what changes in the architecture, the consistency guarantees, the latency and the cost. ```hint Before or during the request For each freshness level, decide whether the value is computed before the request arrives or while it is being served, and what sits between the computation and the model. ``` ```hint One definition, two paths If a feature is computed by a streaming job for serving and by a batch job for training, think about what keeps the two from drifting apart. ``` #### What This Part Should Cover - Precomputed batch features in an online store, streaming aggregates, and on-demand features computed at request time - Latency, cost and operational complexity of each option - Keeping the values used in training consistent with the values served online - Staleness, missing values and fallbacks when the fresh path fails ### What a Strong Answer Covers - Concrete, prioritized risks specific to feature pipelines rather than code-style comments - Leakage and training-serving skew identified as the most damaging failures, because they are silent - A serving design matched to how much freshness each feature actually needs - A rollout and monitoring plan that would catch silent data problems ### Follow-up Questions - The PR changes the definition of a feature that the deployed model already uses. How do you roll it out safely? - How would you detect training-serving skew in production? - The streaming job that maintains a recent-activity feature falls an hour behind. What should the model see, and how should the system behave? - Which tests would you add to the PR so that the problems you found would have failed CI?

Overview: Review a feature pipeline pull request whose tests all pass and explain what could break once it runs in production, from time leakage and training-serving skew to data quality and scale. Then decide how the same features should be served when the model needs day-old, minute-fresh or request-time values.

Read the full Netflix Machine Learning Engineer interview experience this question came from

|Home/ML System Design/Netflix
Netflix logo
Netflix
Sep 28, 2026
hardMachine Learning EngineerTechnical ScreenML System Design
0
0

In the practical coding half of a machine learning screen, the interviewer pastes a pull request that changes a feature pipeline: the code that turns raw event data into features for a model. Every test in the PR passes. You are asked to review it live and explain what problems the change could cause if it lands in production. The conversation then moves on to how these features should be served when the model needs them at different levels of freshness.

Constraints and Clarifications

  • The PR's code is not reproduced here. Practice by naming what you would check in such a PR, what each problem would look like in the code, and how it would surface in production. If you practice with a partner, have them write a short feature-computation function for you to review.
  • "Tests pass" means the logic is correct on the test fixtures. The question is about everything those fixtures do not exercise.

Clarifying Questions Guidance

  • Are these features used for training, for online inference, or both? If both, is the same code used to compute them offline and online?
  • What are the inputs (event streams, warehouse tables), how large are they, and how late can events arrive?
  • Is the pipeline a scheduled batch job, a streaming job, or code that runs at request time?
  • What happens downstream when a feature value is missing or stale?

Part 1 — Production risks in a PR whose tests pass

Review the PR. What could go wrong once it runs in production, even though its tests pass? Rank the problems by impact, and for each one, say how you would catch it before release or detect it afterwards.

What This Part Should Cover Guidance

  • Correctness at each point in time: leakage of future data, late and out-of-order events, timezone and window boundaries
  • Data quality: nulls, duplicates, join fan-out, and schema or unit changes upstream
  • Consistency between training and serving, and the rollout of a changed feature definition
  • Operational risks (skewed keys, memory, runtime, idempotent reruns and backfills) and the checks and monitoring that catch them

Part 2 — Serving the features at different freshness requirements

How would you serve these features if the model can use values that are a day old, if they must reflect events from the last few minutes, or if they depend on the request itself? Explain what changes in the architecture, the consistency guarantees, the latency and the cost.

What This Part Should Cover Guidance

  • Precomputed batch features in an online store, streaming aggregates, and on-demand features computed at request time
  • Latency, cost and operational complexity of each option
  • Keeping the values used in training consistent with the values served online
  • Staleness, missing values and fallbacks when the fresh path fails

What a Strong Answer Covers Guidance

  • Concrete, prioritized risks specific to feature pipelines rather than code-style comments
  • Leakage and training-serving skew identified as the most damaging failures, because they are silent
  • A serving design matched to how much freshness each feature actually needs
  • A rollout and monitoring plan that would catch silent data problems

Follow-up Questions Guidance

  • The PR changes the definition of a feature that the deployed model already uses. How do you roll it out safely?
  • How would you detect training-serving skew in production?
  • The streaming job that maintains a recent-activity feature falls an hour behind. What should the model see, and how should the system behave?
  • Which tests would you add to the PR so that the problems you found would have failed CI?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...