Classify Variable-Length, Unordered Point Sequences with Missing Features
Company: Citadel
Role: Data Scientist
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
You are in a live coding round with a plain Python environment in which PyTorch is available. Within about 35 minutes of working time you must end up with a model that trains and runs inference end to end. The problem is revealed in stages, and each stage removes an assumption the previous one relied on.
The data is a collection of labeled examples. Each example is a sequence of points, each point has 6 numeric features, and each whole sequence carries a binary label (0 or 1). The goal is to predict the label of a whole sequence.
### Clarifying Questions
- Roughly how many labeled sequences are there, and how balanced are the two classes?
- Are the 6 features continuous and on comparable scales, or do some need normalization?
- Which libraries may be used besides PyTorch (for example numpy or scikit-learn), and does "build the model" mean writing the training loop yourself?
- How is the result judged at the end: a held-out metric such as accuracy or ROC AUC, or only that training and inference run?
### Part 1 — Fixed-length time series
Every sequence has exactly 16 points with 6 features each, and the points are in time order (the interviewer confirms this is a time series). Build a binary classifier for the whole sequence, train it, and run inference with it.
```hint Size the model to the input
A single example is only 16 × 6 numbers. Before choosing an architecture, ask what the smallest model is that can use all of them and have a training loop running within minutes.
```
#### What This Part Should Cover
- A concrete input tensor shape and a model that maps it to one logit per sequence.
- A complete training loop and an inference function that actually run.
- A justified choice between a simple baseline and a sequence architecture, given the time budget.
### Part 2 — Variable sequence length
Sequences now contain anywhere from 16 to 100 points. You may not resample or augment the data (no upsampling or downsampling to a fixed length); the interviewer wants the variable length handled inside the model.
```hint Fixed size out of variable size in
Think about which operations turn any number of per-point vectors into one vector of fixed size, and how padded positions must be kept from influencing them.
```
#### Clarifying Questions for this Part
- Could the length of a sequence itself be predictive of the label?
#### What This Part Should Cover
- A batching scheme for unequal lengths (padding with a mask, or packing).
- A length-independent reduction and why it stays correct in the presence of padding.
- Why resampling is ruled out and what the model-based alternative preserves.
### Part 3 — Point order is not fixed
The points inside a sequence no longer arrive in a meaningful order. The interviewer asks whether a linear model could handle this. Adapt your approach, then train and run inference with the adapted model before time runs out.
```hint Test with a shuffle
If shuffling the points of an example must never change its prediction, go through each component of your current model and ask which ones would produce a different output after a shuffle.
```
#### Clarifying Questions for this Part
- Does "order not fixed" mean the order carries no information at all, or that time information is available in some other form?
#### What This Part Should Cover
- Recognizing that each example is now a set, and which earlier choices (positional encodings, recurrent or convolutional layers over time) no longer fit.
- A permutation-invariant representation and a direct answer to the linear-model question.
- A simple, fast-to-train model on that representation, with its training and inference path.
### Part 4 — Missing feature values
Some feature values are missing (NaN). How does your model handle them?
```hint Compare the options
Work out what filling with the mean, filling with zero, and leaving the value missing each do to your specific model, and whether the fact that a value is missing could itself be informative.
```
#### What This Part Should Cover
- How missing values flow through feature construction and the model without crashing or silently producing a NaN loss.
- The effect of mean filling versus zero filling versus native missing-value handling for the chosen model family.
- Preserving missingness as a signal, and avoiding leakage when computing fill statistics.
### What a Strong Answer Covers
- Starting from the simplest model that respects the structure of the data instead of defaulting to a heavy architecture.
- Explicit reasoning about invariances: to length in Part 2 and to order in Part 3.
- A model that trains and predicts end to end within the time box, updated at every stage.
- Missing-value handling matched to the model family, and a sound evaluation plan.
### Follow-up Questions
- How would you choose between hand-built summary statistics with a tree and a learned per-point encoder with pooling, and what evidence would make you switch?
- If the label depended on relationships between pairs of points, how would a permutation-invariant model capture that?
- If a single decision tree underfits, how would you move to an ensemble of trees, and what changes in how missing values are handled?
- If the labeled dataset is small, how would you detect and control overfitting in these models?
Overview: A staged live-coding ML task: classify whole sequences of 6-feature points as 0 or 1, then adapt the model as lengths vary from 16 to 100, point order stops being meaningful, and some features go missing. Tests model selection, length and permutation invariance, and practical missing-value handling.