Explain Random Forest randomness and implications
Company: Snapchat
Role: Data Scientist
Category: Machine Learning
Difficulty: hard
Interview Round: Onsite
Random Forest rigor: 1) Enumerate all sources of randomness (bootstrap sampling, feature subsampling at each split, random tie-breaking, randomized split points) and explain the effect of each on bias and variance. 2) For a dataset with 100k rows, 100 features, and a 5% positive rate, propose n_estimators, max_depth, and max_features; justify how max_features controls tree correlation. 3) Compare out-of-bag (OOB) error to 5-fold cross-validation; when can they disagree and why? 4) Why are impurity-based importances biased toward continuous or high-cardinality features? Propose and justify a corrected approach (e.g., permutation importance with stratified shuffles and repeated runs). 5) Outline strategies for class imbalance (class_weight, threshold moving, balanced subsampling) and discuss consequences for probability calibration and decision thresholds.
Overview: This question evaluates a candidate's understanding of Random Forest ensemble mechanics—sources of randomness, their effects on bias and variance, hyperparameter impacts, evaluation choices (OOB vs cross-validation), feature-importance bias, and class-imbalance strategies—within the Machine Learning domain for binary classification.
Read the full Snapchat Data Scientist interview experience this question came from