Build a Heart Disease Baseline

Read the full interview experience this question came from →

Quick Overview

This question evaluates proficiency in applied machine learning workflows, including data inspection and cleaning, exploratory data analysis, feature engineering, baseline binary classification modeling, and evaluation using Python libraries like pandas and seaborn.

Build a Heart Disease Baseline

Company: Hudson River Trading

Role: Software Engineer

Category: Machine Learning

Difficulty: hard

Interview Round: Onsite

You are given a tabular dataset for predicting whether a patient has heart disease. The dataset contains a binary target column such as `has_heart_disease` and several features, for example age, height, weight, blood pressure, cholesterol, smoking status, and other clinical measurements. Using Python, `pandas`, and `seaborn`, walk through how you would: 1. Load and inspect the data. 2. Clean missing values, duplicates, and obviously invalid records. 3. Perform exploratory data analysis and visualize relationships between features and the target. 4. Engineer useful features when appropriate, such as BMI from height and weight. 5. Train a reasonable baseline model for binary classification. 6. Evaluate the model and explain which metrics you would report. 7. Summarize what patterns you found and what you would check before using the model in practice. You may assume standard Python ML libraries are available.

Overview: This question evaluates proficiency in applied machine learning workflows, including data inspection and cleaning, exploratory data analysis, feature engineering, baseline binary classification modeling, and evaluation using Python libraries like pandas and seaborn.

Read the full Hudson River Trading Software Engineer interview experience this question came from

|Home/Machine Learning/Hudson River Trading
Hudson River Trading logo
Hudson River Trading
Mar 13, 2026
hardSoftware EngineerOnsiteMachine Learning
9
0

You are given a tabular dataset for predicting whether a patient has heart disease. The dataset contains a binary target column such as has_heart_disease and several features, for example age, height, weight, blood pressure, cholesterol, smoking status, and other clinical measurements.

Using Python, pandas, and seaborn, walk through how you would:

  1. Load and inspect the data.
  2. Clean missing values, duplicates, and obviously invalid records.
  3. Perform exploratory data analysis and visualize relationships between features and the target.
  4. Engineer useful features when appropriate, such as BMI from height and weight.
  5. Train a reasonable baseline model for binary classification.
  6. Evaluate the model and explain which metrics you would report.
  7. Summarize what patterns you found and what you would check before using the model in practice.

You may assume standard Python ML libraries are available.

Loading comments...