Design MapReduce for schedule aggregation

Quick Overview

Design MapReduce for schedule aggregation evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

Design MapReduce for schedule aggregation

Company: Uber

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a MapReduce pipeline that processes large-scale scheduling data (users’ busy intervals) to compute common available time slots of at least duration d for specified groups. Define the map outputs (keys/values), partitioning, and reduce logic; explain how you discretize time, mitigate data skew, and validate results at scale.

Quick Answer: Design MapReduce for schedule aggregation evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

|Home/System Design/Uber
Uber logo
Uber
Aug 1, 2025, 12:00 AM
mediumSoftware EngineerOnsiteSystem Design
6
0

Design MapReduce for schedule aggregation

MapReduce Design: Common Availability From Busy Intervals

Context

You are given large-scale calendar data: each user has 0 or more busy intervals during the day. For a set of specified groups (each group is a set of user IDs), compute the common available time slots whose duration is at least d minutes. Assume all timestamps are normalized to UTC and intervals are half-open [start, end).

Input datasets:

  • Busy intervals: records (user_id, start_ts, end_ts)
  • Group membership: records (group_id, user_id)
  • Query parameter: minimum duration d (minutes)

Output:

  • For each (group_id, calendar_day), the list of common free intervals of length ≥ d.

Requirements

  1. Define the Map outputs (keys/values), partitioning, and Reduce logic.
  2. Explain how you discretize time (or avoid discretization) and the trade-offs.
  3. Describe how you mitigate data skew (e.g., very large groups, rush-hour hotspots).
  4. Explain how you validate correctness and performance at scale.

Clarifying Questions to Ask Guidance

  • Clarify users, core use cases, read/write patterns, scale, latency, availability, and data retention.
  • State explicit assumptions before making sizing or architecture decisions.
  • Prioritize the functional path first, then address reliability, security, observability, and rollout.

What a Strong Answer Covers Guidance

  • A scoped requirements summary with concrete non-goals and success metrics.
  • API, data model, architecture, consistency, capacity, and operations.
  • Reasoned trade-offs among simple and scalable designs, including bottlenecks and failure modes.
  • A validation, monitoring, migration, and launch plan appropriate for the risk level.

Follow-up Questions Guidance

  • What breaks first at 10x traffic or data volume?
  • How would you degrade gracefully during dependency failures?
  • What metrics and alerts would prove the design is healthy after launch?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...