Predict Employee Attrition from an HR Dataset and Explain Its Drivers

Read the full interview experience this question came from →

Quick Overview

A machine learning design exercise: given a tab-separated HR dataset with features such as age, job role, monthly income and overtime, predict which employees will leave and explain which factors drive attrition. It covers feature processing, model choice, class imbalance, evaluation metrics, threshold selection and interpretability.

Predict Employee Attrition from an HR Dataset and Explain Its Drivers

Company: Microsoft

Role: Applied Scientist

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

You are given an HR dataset, `attritiondata.tsv`: a tab-separated file with one row per employee, feature columns such as `Age`, `JobRole`, `MonthlyIncome` and `OverTime`, and a label recording whether the employee left the company. Build a model that predicts employee attrition, and use it to analyze which factors are associated with employees leaving. The exercise runs in an online coding environment. You write the data-processing code and choose and train a model, but much of the evaluation is the design discussion around feature processing, model selection, class imbalance, evaluation metrics and interpretability. Expect follow-up questions on the machine learning behind the model you choose. ### Constraints and Clarifications - Assume the label column is named `Attrition` and holds `Yes` or `No`; adapt if the file differs. - Assume the file fits in memory and that Python with pandas, scikit-learn and a gradient-boosting library is available. - The exact columns beyond the four named above, the number of rows and the share of leavers are found by inspecting the file. ### Clarifying Questions - What share of the employees in the file left? - Is the file a single snapshot, or does it cover several periods, possibly with several rows per employee? - How will the predictions be used, for example to pick employees for retention conversations, and how many people can HR act on? - Were any columns recorded after an employee resigned, or derived from the outcome? - May attributes such as age be used as inputs to a model that informs decisions about employees? ### Part 1 — Load and prepare the data Load the file and write the preprocessing: decide how each kind of column is represented, what to drop, and how to handle missing and unseen values. ```hint Read the columns before encoding Check each column's type, its number of distinct values and how it relates to the outcome before choosing an encoding, and look for columns that would not exist at the time a prediction is made. ``` #### What This Part Should Cover - An encoding for each kind of column (numeric, yes/no, nominal such as `JobRole`) and the handling of missing and unseen values - Removal of identifiers, constant columns and leakage - Preprocessing fit on training data only and packaged together with the model ### Part 2 — Choose a model and handle class imbalance Pick and train a model, justify it against at least one alternative, and explain how you deal with leavers being the minority class. ```hint Small, tabular and skewed Weigh what each candidate model offers on a modest tabular dataset, and list more than one way to make the minority class count, along with where each one can leak information. ``` #### What This Part Should Cover - A simple baseline and a stronger model, with reasons tied to this data - At least two ways to handle the imbalance, and their trade-offs - The hyperparameters that matter and how to tune them without touching the test data ### Part 3 — Evaluate the model Choose the metrics, the validation scheme and the decision threshold. ```hint Why not accuracy Work out the accuracy of a model that predicts "stays" for everyone, then pick metrics that such a model cannot game. ``` #### What This Part Should Cover - Metrics suited to an imbalanced label, and why accuracy misleads - A validation scheme that matches how the data was collected - A threshold chosen on validation data from the cost of errors or from HR's capacity ### Part 4 — Explain the attrition factors Explain which factors the model associates with attrition, across all employees and for an individual employee, and how HR should read the results. ```hint Importance is not cause Think about what a feature-importance number actually measures, how correlated features such as income and job role share it, and which claims it does and does not support. ``` #### What This Part Should Cover - Global and per-employee explanations suited to the chosen model - The effect of correlated features, and the stability of the explanations - The limits of the conclusions: association rather than causation, and sensitive attributes ### What a Strong Answer Covers - A leak-free pipeline from the raw file to a prediction, with preprocessing inside cross-validation - A model choice justified by the data's size, column types and imbalance, compared with a baseline - Metrics, validation and a threshold that reflect the minority class and how predictions will be used - Explanations of the attrition drivers, with their caveats - Working, readable code for the data processing and training ### Follow-up Questions - If you chose a gradient-boosted tree model, how does it differ from a random forest, and how do the learning rate, tree depth and number of trees interact? - If you reweight the classes, what happens to the predicted probabilities, and how would you fix them if HR wants a calibrated risk score? - The model scores employees every month. How would you detect data drift, and when would you retrain? - The explanation ranks overtime as the top factor. Can HR conclude that reducing overtime will reduce attrition?

Overview: A machine learning design exercise: given a tab-separated HR dataset with features such as age, job role, monthly income and overtime, predict which employees will leave and explain which factors drive attrition. It covers feature processing, model choice, class imbalance, evaluation metrics, threshold selection and interpretability.

Read the full Microsoft Applied Scientist interview experience this question came from

|Home/ML System Design/Microsoft
Microsoft logo
Microsoft
Aug 27, 2026
mediumApplied ScientistOnsiteML System Design
0
0

You are given an HR dataset, attritiondata.tsv: a tab-separated file with one row per employee, feature columns such as Age, JobRole, MonthlyIncome and OverTime, and a label recording whether the employee left the company. Build a model that predicts employee attrition, and use it to analyze which factors are associated with employees leaving.

The exercise runs in an online coding environment. You write the data-processing code and choose and train a model, but much of the evaluation is the design discussion around feature processing, model selection, class imbalance, evaluation metrics and interpretability. Expect follow-up questions on the machine learning behind the model you choose.

Constraints and Clarifications

  • Assume the label column is named Attrition and holds Yes or No ; adapt if the file differs.
  • Assume the file fits in memory and that Python with pandas, scikit-learn and a gradient-boosting library is available.
  • The exact columns beyond the four named above, the number of rows and the share of leavers are found by inspecting the file.

Clarifying Questions Guidance

  • What share of the employees in the file left?
  • Is the file a single snapshot, or does it cover several periods, possibly with several rows per employee?
  • How will the predictions be used, for example to pick employees for retention conversations, and how many people can HR act on?
  • Were any columns recorded after an employee resigned, or derived from the outcome?
  • May attributes such as age be used as inputs to a model that informs decisions about employees?

Part 1 — Load and prepare the data

Load the file and write the preprocessing: decide how each kind of column is represented, what to drop, and how to handle missing and unseen values.

What This Part Should Cover Guidance

  • An encoding for each kind of column (numeric, yes/no, nominal such as JobRole ) and the handling of missing and unseen values
  • Removal of identifiers, constant columns and leakage
  • Preprocessing fit on training data only and packaged together with the model

Part 2 — Choose a model and handle class imbalance

Pick and train a model, justify it against at least one alternative, and explain how you deal with leavers being the minority class.

What This Part Should Cover Guidance

  • A simple baseline and a stronger model, with reasons tied to this data
  • At least two ways to handle the imbalance, and their trade-offs
  • The hyperparameters that matter and how to tune them without touching the test data

Part 3 — Evaluate the model

Choose the metrics, the validation scheme and the decision threshold.

What This Part Should Cover Guidance

  • Metrics suited to an imbalanced label, and why accuracy misleads
  • A validation scheme that matches how the data was collected
  • A threshold chosen on validation data from the cost of errors or from HR's capacity

Part 4 — Explain the attrition factors

Explain which factors the model associates with attrition, across all employees and for an individual employee, and how HR should read the results.

What This Part Should Cover Guidance

  • Global and per-employee explanations suited to the chosen model
  • The effect of correlated features, and the stability of the explanations
  • The limits of the conclusions: association rather than causation, and sensitive attributes

What a Strong Answer Covers Guidance

  • A leak-free pipeline from the raw file to a prediction, with preprocessing inside cross-validation
  • A model choice justified by the data's size, column types and imbalance, compared with a baseline
  • Metrics, validation and a threshold that reflect the minority class and how predictions will be used
  • Explanations of the attrition drivers, with their caveats
  • Working, readable code for the data processing and training

Follow-up Questions Guidance

  • If you chose a gradient-boosted tree model, how does it differ from a random forest, and how do the learning rate, tree depth and number of trees interact?
  • If you reweight the classes, what happens to the predicted probabilities, and how would you fix them if HR wants a calibrated risk score?
  • The model scores employees every month. How would you detect data drift, and when would you retrain?
  • The explanation ranks overtime as the top factor. Can HR conclude that reducing overtime will reduce attrition?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...