Build DiD dataset with SQL
Company: Amazon
Role: Data Scientist
Category: Data Manipulation (SQL/Python)
Difficulty: medium
Interview Round: Technical Screen
Using the schema and sample data below, write SQL to build an individual-day panel suitable for staggered-adoption DiD of the shuttle’s effect on participation. Requirements: A) output columns: employee_id, site_id, date, participated, adoption_date (per site), treated_site (1 if date >= adoption_date at that site, else 0; never-treated have NULL adoption_date and 0), event_time_days = date - adoption_date (NULL for never-treated), and a binary post indicator; B) ensure no off-by-one errors on the adoption boundary; C) also produce a weekly site-level table with participation_rate = avg(participated) per site-week, correctly handling sites without shuttle; D) assume employees do not move sites. Provide SQL that works on a modern warehouse (e.g., BigQuery or PostgreSQL). Schema:
sites(site_id, city)
employees(employee_id, site_id, hire_date)
shuttle_service(site_id, start_date) -- present only for treated sites
participation(employee_id, date, participated)
Sample tables:
sites
+---------+------+
| site_id | city |
+---------+------+
| 1 | SEA |
| 2 | NYC |
+---------+------+
employees
+-------------+---------+------------+
| employee_id | site_id | hire_date |
+-------------+---------+------------+
| 101 | 1 | 2024-10-01 |
| 102 | 1 | 2025-01-01 |
| 201 | 2 | 2024-11-15 |
+-------------+---------+------------+
shuttle_service
+---------+------------+
| site_id | start_date |
+---------+------------+
| 1 | 2025-01-15 |
+---------+------------+
participation
+-------------+------------+--------------+
| employee_id | date | participated |
+-------------+------------+--------------+
| 101 | 2025-01-10 | 1 |
| 101 | 2025-01-20 | 1 |
| 102 | 2025-01-20 | 0 |
| 201 | 2025-01-10 | 1 |
| 201 | 2025-01-20 | 0 |
+-------------+------------+--------------+
Overview: This question evaluates proficiency in SQL-based data manipulation and panel construction for staggered-adoption difference-in-differences, testing skills in handling adoption dates, event-time calculations, boundary conditions, and aggregation in the Data Manipulation (SQL/Python) domain.
Read the full Amazon Data Scientist interview experience this question came from
Build an individual-day DiD panel with staggered site adoption
Using the tables below, write SQL to build an individual-day panel suitable for a staggered-adoption Difference-in-Differences analysis of a shuttle service’s effect on participation.
Output requirements (one row per employee per date):
1) Columns: employee_id, site_id, date, participated, adoption_date (per site), treated_site (1 if date >= adoption_date for that site, else 0; never-treated sites must have NULL adoption_date and treated_site = 0), event_time_days = date - adoption_date (NULL for never-treated), and post (binary indicator equal to 1 if date >= adoption_date for treated sites, else 0).
2) Avoid off-by-one errors: the adoption boundary is inclusive (i.e., date = adoption_date is treated/post).
3) Assume employees never change sites.
Panel construction rules:
- Use the distinct dates present in the participation table as the date spine.
- Include all employees for all spine dates on/after their hire_date.
- If an employee has no participation row for a given spine date, set participated = 0.
Tables
sites(site_id INT, city VARCHAR(10))
employees(employee_id INT, site_id INT, hire_date DATE)
shuttle_service(site_id INT, start_date DATE)
participation(employee_id INT, date DATE, participated INT)
Hints
- Create a date spine (e.g., distinct dates from participation) then CROSS JOIN employees to it to form the panel.
- LEFT JOIN shuttle_service at the site level to get adoption_date; use date >= adoption_date (inclusive) to avoid off-by-one errors.
Aggregate to a weekly site-level participation rate (including never-treated sites)
Using the same schema, write SQL to produce a weekly site-level table with:
- site_id
- week_start_date (week bucket start; you may use DATE_TRUNC('week', date) in PostgreSQL)
- participation_rate = AVG(participated)
Requirements:
1) Correctly handle sites without shuttle service (they must still appear for weeks where they have employee-day rows in the panel).
2) Use the same paneling rule as in the individual-day dataset: for each site, include all employee-day rows formed by (employees at the site) x (distinct participation dates) on/after hire_date, with missing participation treated as 0.
Return one row per site-week.
Tables
sites(site_id INT, city VARCHAR(10))
employees(employee_id INT, site_id INT, hire_date DATE)
shuttle_service(site_id INT, start_date DATE)
participation(employee_id INT, date DATE, participated INT)
Hints
- Build (or reuse) the employee-day panel first, then aggregate by site_id and DATE_TRUNC('week', date).
- To ensure never-treated sites are included, avoid inner joins to shuttle_service when creating the weekly table.