Design a Multi-Region CI/CD System with Separate Control and Data Planes
Company: Snapchat
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a CI/CD system in which the control plane is separated from the data plane. The system takes code changes through build and test, produces release artifacts, and deploys them to a fleet of hosts spread across multiple regions. Distributing work and releases to hosts at the region level is an explicit requirement.
### Constraints and Clarifications
- The problem was posed at this level: CI/CD, a separated control plane and data plane, and region-level distribution to hosts. Fleet size, number of regions, deployment frequency and reliability targets were not given; ask for them or state your assumptions.
- Focus on the architecture of the CI/CD system itself, not on the services it deploys.
### Clarifying Questions
- Does "distribution" mean placing build and test jobs on runner hosts in each region, rolling releases out to production hosts region by region, or both?
- Are the target hosts long-lived machines (virtual machines or bare metal) or containers scheduled by an orchestrator?
- How many regions and hosts are there, and how often are releases deployed?
- If a region loses contact with the control plane, must its hosts keep serving their current version, and may the region keep deploying?
- Who defines the rollout policy: each service's pipeline configuration, or a central policy?
### Part 1 — Separate the control plane from the data plane
Lay out the components of the system and decide which belong to the control plane and which to the data plane. Define the interface between the two planes, and walk through one change from commit to running on hosts.
```hint Deciding versus doing
Sort each component by whether it decides what should happen or does the work on a host, then check how often each side must talk to the other and what breaks when that conversation stops.
```
#### What This Part Should Cover
- Control-plane responsibilities: pipeline definitions, desired state, scheduling and rollout orchestration, audit
- Data-plane responsibilities: build runners, artifact storage and transfer, host agents that apply releases
- The contract between the planes (desired versus observed state, push versus pull), and how it is secured
- The end-to-end path of one change
### Part 2 — Region-level distribution to hosts
Design how work and releases are distributed across regions and to the hosts within each region. Cover how a release progresses from region to region, how artifacts reach the hosts in each region, and how a region stays safe when something goes wrong in the middle of a rollout.
```hint Blast radius
Decide the smallest set of hosts a bad release may reach before the system notices, and how that set is allowed to grow.
```
```hint Bytes travel far only once
Count how many times the same artifact would cross region boundaries if every host fetched it from a single origin.
```
#### What This Part Should Cover
- Region-scoped components and how hosts are grouped within a region
- Wave-based rollout across and within regions, with health gates and automatic halt or rollback
- Artifact replication per region and efficient fan-out to hosts
- Behavior when the control plane is lost, when a region is partitioned, and when part of a region fails
### What a Strong Answer Covers
- A clear split in which the data plane keeps working while the control plane is down
- Declarative desired state reconciled by agents, with idempotent operations
- Regional isolation and staged rollouts that limit the blast radius of a bad release
- Scalable artifact distribution and host status reporting
- Safety: artifact verification, rollback, access control and audit
- Observability of rollout progress and fleet convergence
### Follow-up Questions
- The control plane is down for an hour in the middle of a multi-region rollout. What happens to hosts already updated, hosts not yet updated, and the next deploy?
- A release passes its canary in the first region but fails in the third. How does the system detect it, stop, and roll back, and what happens in the regions already done?
- How would you ship an emergency fix to every region quickly without abandoning the safety gates?
- How do you stop a compromised build runner from shipping a tampered artifact to production hosts?
Overview: System design question about a CI/CD platform whose control plane is separated from its data plane. Candidates define what each plane owns and how the planes communicate, then design region-level distribution of builds and releases to hosts, including staged rollouts, artifact replication and failure handling.
Read the full Snapchat Software Engineer interview experience this question came from