Quick Overview

This question evaluates data manipulation and exploratory data analysis skills, including aggregation, outlier detection (IQR/Z-score) and visualization using Python libraries such as pandas and matplotlib/seaborn.

Visualize and Clean SKU Sales Data for Outliers

Company: Boston Consulting Group

Role: Data Scientist

Category: Data Manipulation (SQL/Python)

Difficulty: medium

Interview Round: Technical Screen

sales_data +------------+--------+-----------+----------+------------+---------+ | date | sku_id | unit_sold | revenue | promo_flag | store_id| +------------+--------+-----------+----------+------------+---------+ | 2023-01-01 | A123 | 120 | 2400.00 | 1 | S01 | | 2023-01-02 | A123 | 80 | 1600.00 | 0 | S01 | | 2023-01-01 | B456 | 200 | 3000.00 | 1 | S02 | | 2023-01-02 | B456 | 50 | 750.00 | 0 | S02 | ##### Scenario Codesignal live-coding: analyst receives raw daily SKU sales data and must explore it visually while cleaning extreme values. ##### Question Using Python (pandas, matplotlib/seaborn), draw a histogram of daily revenue per SKU and identify outliers with the IQR or Z-score method. Remove the detected outliers and re-plot the cleaned distribution. ##### Hints Focus on reproducible pandas pipeline: load → aggregate → detect outliers → filter → visualize before/after.

Overview: This question evaluates data manipulation and exploratory data analysis skills, including aggregation, outlier detection (IQR/Z-score) and visualization using Python libraries such as pandas and matplotlib/seaborn.

You are given raw daily SKU sales data at the transaction level. Write a SQL query that: 1) Aggregates revenue to the level of (date, sku_id) as daily_revenue. 2) For each sku_id, computes the 25th percentile (Q1) and 75th percentile (Q3) of daily_revenue using percentile_cont. 3) Derives the IQR-based outlier bounds: lower_bound = Q1 - 1.5 * (Q3 - Q1) and upper_bound = Q3 + 1.5 * (Q3 - Q1). 4) Flags each (date, sku_id) daily_revenue as an outlier if it falls outside these bounds. Return one row per (date, sku_id) with the daily_revenue, Q1, Q3, lower_bound, upper_bound, and a boolean is_outlier flag.

Tables

sales_data(date DATE, sku_id VARCHAR(10), unit_sold INTEGER, revenue DECIMAL(12,2), promo_flag INTEGER, store_id VARCHAR(10))

Hints

  1. First aggregate to daily revenue per SKU: SELECT date, sku_id, SUM(revenue) AS daily_revenue FROM sales_data GROUP BY date, sku_id.
  2. Use percentile_cont(0.25) and percentile_cont(0.75) WITHIN GROUP (ORDER BY daily_revenue) over each sku_id to compute Q1 and Q3.

Loading coding console...