Design a Feature Store for Offline Training and Low-Latency Online Serving
Company: Unity
Role: Software Engineer
Category: ML System Design
Difficulty: medium
Interview Round: Onsite
Design a feature store that supports both offline machine learning training workflows and online low-latency serving. It must handle high-volume event data, give point-in-time correct retrieval of features for training, and serve features at low latency in production.
### Constraints and Clarifications
- Event volume, the number of entities and features, freshness requirements and the serving latency target are not given. Ask for them or state your assumptions.
- Point-in-time correct retrieval means that a training example whose label is observed at time T receives only the feature values that were known at time T.
- Two kinds of clients use the system: data scientists building training sets, and production model servers requesting features for live predictions.
### Clarifying Questions
- Which entities are features keyed by (for example users, items or sessions), and how many of each exist?
- What event rate must ingestion sustain, on average and at peak?
- How fresh must online features be: seconds after an event, or is an hourly or daily refresh acceptable for some of them?
- What latency must an online lookup meet, at which percentile and request rate, and how many entities and features does one request need?
- Who defines feature transformations, and in what language? Must one definition feed both training and serving?
- How far back must training data reach, and how long must feature history be kept?
### Part 1 — Ingestion, computation and storage
Design how high-volume event data enters the system, how features are defined and computed from it, and where feature values are stored for training and for serving.
```hint One definition, two consumers
Training needs every historical value of a feature; serving needs only the latest one. Think about whether one store can serve both access patterns well, and how to keep the two views from drifting apart.
```
```hint Not every feature needs seconds
Some features summarize the last few minutes, others the last month. Consider whether they should be computed by the same pipeline.
```
#### What This Part Should Cover
- Feature definitions, versioning and a registry shared by every pipeline
- Streaming and batch computation paths, including backfills
- An offline store for full history and an online store for latest values, with their layouts
- How online and offline values are kept consistent
### Part 2 — Point-in-time correct training data
Given a table of training examples, each with entity keys, a timestamp and a label, produce a training set that attaches to each example the feature values as they were known at the example's timestamp.
```hint What a plain join attaches
Imagine joining last month's examples to the feature table on entity key alone. Which feature value would each example receive?
```
```hint True versus known
A value can describe the world at one time but reach the feature store later. Decide which of those two times the training join must respect.
```
#### What This Part Should Cover
- Exact as-of join semantics, including availability delay and feature expiry
- An implementation that scales to large histories
- Reproducibility of a training set
- Checks that show no future information leaked
### Part 3 — Low-latency online serving
Design the online read path that serves features to production models, and explain how it keeps latency low and how it behaves when parts of it fail.
```hint Count the reads
Work out how many store reads one prediction request triggers when it scores many entities with features from several groups, and how that number drives tail latency.
```
```hint Plan for the miss
Decide what the model receives when a value is missing, stale, or the store does not answer in time.
```
#### What This Part Should Cover
- The serving API and a key layout that keeps reads per request low
- Freshness from streaming writes, and idempotent updates
- Tail-latency control under fan-out, hot keys and growth
- Fallbacks for missing values, stale values and outages
### What a Strong Answer Covers
- Requirements and estimates that drive the choice of stores and pipelines
- A write path that computes each feature once and feeds both stores
- Leakage-free training data, including the difference between event time and availability time
- A read path with a defended tail-latency target and explicit failure behavior
- Detection of training-serving skew, plus freshness and data-quality monitoring
### Follow-up Questions
- A data scientist changes a feature's definition. How do you backfill its history and roll it out without training-serving skew?
- Events arrive hours late. How does that affect the online value and point-in-time correctness?
- How would you detect training-serving skew in production?
- One entity receives a large share of events and lookups. What breaks, and how do you handle it?
Overview: Design a feature store that ingests high-volume event data, produces point-in-time correct training sets, and serves the latest features at low latency in production. It tests streaming and batch feature computation, leakage-free historical joins, online store design, and keeping training and serving consistent.
Read the full Unity Software Engineer interview experience this question came from