Diagnose and fix underperforming ML model
Company: Amazon
Role: Data Scientist
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
You inherited a binary fraud model with extreme class imbalance (positives ≈2%). Current performance on a temporally separated validation set: AUC=0.61, precision@recall=0.90 is only 0.05. You have one day to meaningfully improve recall at fixed review capacity. 1) Describe how you would quickly diagnose underfitting vs. overfitting (learning curves, calibration plots, PR vs. ROC trade-offs, leakage checks). 2) Propose three targeted interventions that can be implemented in a day (e.g., class-weighted loss, monotonic gradient boosting with categorical encoders, threshold moving with cost-sensitive utility) and justify why each should help. 3) Show how you would choose a decision threshold that maximizes expected utility given: FP cost=$2, FN cost=$50, review capacity=0.5% of traffic; write the utility formula and outline the validation-time procedure. 4) List the minimal logging/monitoring you’d add at deployment to detect drift and data quality issues within a week.
Overview: This question evaluates a data scientist's competency in diagnosing and remediating underperforming binary classifiers under severe class imbalance, covering validation diagnostics, calibration, threshold selection under operational review constraints, cost-sensitive utility reasoning, and basic deployment monitoring for drift.
Read the full Amazon Data Scientist interview experience this question came from