Quick Overview

This question evaluates proficiency in SQL-based data manipulation and panel construction for staggered-adoption difference-in-differences, testing skills in handling adoption dates, event-time calculations, boundary conditions, and aggregation in the Data Manipulation (SQL/Python) domain.

Build DiD dataset with SQL

Company: Amazon

Role: Data Scientist

Category: Data Manipulation (SQL/Python)

Difficulty: medium

Interview Round: Technical Screen

Using the schema and sample data below, write SQL to build an individual-day panel suitable for staggered-adoption DiD of the shuttle’s effect on participation. Requirements: A) output columns: employee_id, site_id, date, participated, adoption_date (per site), treated_site (1 if date >= adoption_date at that site, else 0; never-treated have NULL adoption_date and 0), event_time_days = date - adoption_date (NULL for never-treated), and a binary post indicator; B) ensure no off-by-one errors on the adoption boundary; C) also produce a weekly site-level table with participation_rate = avg(participated) per site-week, correctly handling sites without shuttle; D) assume employees do not move sites. Provide SQL that works on a modern warehouse (e.g., BigQuery or PostgreSQL). Schema: sites(site_id, city) employees(employee_id, site_id, hire_date) shuttle_service(site_id, start_date) -- present only for treated sites participation(employee_id, date, participated) Sample tables: sites +---------+------+ | site_id | city | +---------+------+ | 1 | SEA | | 2 | NYC | +---------+------+ employees +-------------+---------+------------+ | employee_id | site_id | hire_date | +-------------+---------+------------+ | 101 | 1 | 2024-10-01 | | 102 | 1 | 2025-01-01 | | 201 | 2 | 2024-11-15 | +-------------+---------+------------+ shuttle_service +---------+------------+ | site_id | start_date | +---------+------------+ | 1 | 2025-01-15 | +---------+------------+ participation +-------------+------------+--------------+ | employee_id | date | participated | +-------------+------------+--------------+ | 101 | 2025-01-10 | 1 | | 101 | 2025-01-20 | 1 | | 102 | 2025-01-20 | 0 | | 201 | 2025-01-10 | 1 | | 201 | 2025-01-20 | 0 | +-------------+------------+--------------+

Overview: This question evaluates proficiency in SQL-based data manipulation and panel construction for staggered-adoption difference-in-differences, testing skills in handling adoption dates, event-time calculations, boundary conditions, and aggregation in the Data Manipulation (SQL/Python) domain.

Read the full Amazon Data Scientist interview experience this question came from

Build an individual-day DiD panel with staggered site adoption

Using the tables below, write SQL to build an individual-day panel suitable for a staggered-adoption Difference-in-Differences analysis of a shuttle service’s effect on participation. Output requirements (one row per employee per date): 1) Columns: employee_id, site_id, date, participated, adoption_date (per site), treated_site (1 if date >= adoption_date for that site, else 0; never-treated sites must have NULL adoption_date and treated_site = 0), event_time_days = date - adoption_date (NULL for never-treated), and post (binary indicator equal to 1 if date >= adoption_date for treated sites, else 0). 2) Avoid off-by-one errors: the adoption boundary is inclusive (i.e., date = adoption_date is treated/post). 3) Assume employees never change sites. Panel construction rules: - Use the distinct dates present in the participation table as the date spine. - Include all employees for all spine dates on/after their hire_date. - If an employee has no participation row for a given spine date, set participated = 0.

Tables

sites(site_id INT, city VARCHAR(10))

employees(employee_id INT, site_id INT, hire_date DATE)

shuttle_service(site_id INT, start_date DATE)

participation(employee_id INT, date DATE, participated INT)

Hints

  1. Create a date spine (e.g., distinct dates from participation) then CROSS JOIN employees to it to form the panel.
  2. LEFT JOIN shuttle_service at the site level to get adoption_date; use date >= adoption_date (inclusive) to avoid off-by-one errors.

Aggregate to a weekly site-level participation rate (including never-treated sites)

Using the same schema, write SQL to produce a weekly site-level table with: - site_id - week_start_date (week bucket start; you may use DATE_TRUNC('week', date) in PostgreSQL) - participation_rate = AVG(participated) Requirements: 1) Correctly handle sites without shuttle service (they must still appear for weeks where they have employee-day rows in the panel). 2) Use the same paneling rule as in the individual-day dataset: for each site, include all employee-day rows formed by (employees at the site) x (distinct participation dates) on/after hire_date, with missing participation treated as 0. Return one row per site-week.

Tables

sites(site_id INT, city VARCHAR(10))

employees(employee_id INT, site_id INT, hire_date DATE)

shuttle_service(site_id INT, start_date DATE)

participation(employee_id INT, date DATE, participated INT)

Hints

  1. Build (or reuse) the employee-day panel first, then aggregate by site_id and DATE_TRUNC('week', date).
  2. To ensure never-treated sites are included, avoid inner joins to shuttle_service when creating the weekly table.

Loading coding console...