Pre-process Financial Data for Linear Regression Modeling
Company: Voleon
Role: Data Scientist
Category: Data Manipulation (SQL/Python)
Difficulty: medium
Interview Round: Technical Screen
market_data
+------------+----------+----------+--------+
| date | feature1 | feature2 | target |
+------------+----------+----------+--------+
| 2024-01-02 | 1.23 | 0.34 | 0.05 |
| 2024-01-03 | 1.25 | 0.30 | -0.01 |
| 2024-01-04 | 1.20 | 0.28 | 0.02 |
| 2024-01-05 | 1.18 | 0.27 | -0.03 |
+------------+----------+----------+--------+
##### Scenario
Voleon DS tech round: pre-processing a financial time-series dataset before modeling.
##### Question
Using the table below, write Python/Pandas code to clean nulls, winsorize extreme values at the 1st/99th percentiles, standardize predictors, and create an X, y pair ready for linear regression.
##### Hints
Focus on dataframe operations: dropna, clip, StandardScaler, and separate features/target.
Overview: This question evaluates data preprocessing and feature engineering skills—handling missing values, winsorizing outliers, standardizing predictors, and assembling X/y datasets using Python/Pandas within the Data Manipulation (SQL/Python) domain.
Preprocess `market_data` into standardized predictors and a target value for linear regression.
Drop rows where `date`, `feature1`, `feature2`, or `target` is `NULL`. On the remaining rows, winsorize `feature1` and `feature2` at the 1st and 99th percentiles, then standardize each winsorized predictor using population mean and population standard deviation.
Return one row per remaining date with these columns:
- `date`
- `x_feature1`: standardized winsorized `feature1`, rounded to 5 decimals
- `x_feature2`: standardized winsorized `feature2`, rounded to 5 decimals
- `y`: the original target rounded to 2 decimals
Order by `date`.
Tables
market_data(date DATE, feature1 DECIMAL(10,4), feature2 DECIMAL(10,4), target DECIMAL(10,4))
Hints
- Use percentile_cont to compute the 1st and 99th percentile bounds.
- Clip values with LEAST and GREATEST before computing standardization statistics.