Design a Control Plane That Manages Cloud Clusters and Monitors Host Health
Company: Snowflake
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design a control plane for a cloud platform that runs clusters of hosts (virtual machines or physical servers). The control plane is the management layer. It accepts requests to manage clusters, decides which hosts belong to which cluster, and drives every cluster to the configuration that was requested. It must also continuously monitor the health of every host and act when a host becomes unhealthy.
Assume the control plane supports at least creating, resizing and deleting clusters, and that an existing infrastructure API can provision and terminate hosts on request. The hosts run the actual workloads (the data plane). The control plane manages them but is not on their request path.
```hint Desired versus actual
Decide what the control plane stores as the source of truth for a cluster, and what should happen when an operation that touches many hosts fails halfway through.
```
```hint Silence is ambiguous
A host that stops reporting may be dead, overloaded, or cut off from the monitor by a network partition. Consider how you would tell these cases apart before replacing it.
```
### Constraints and Clarifications
- No scale, detection-time or availability figures were given. State the numbers you assume and size the design from them.
- Designing the infrastructure API is out of scope. Your control plane calls it, and those calls can fail, time out, or succeed without the response reaching you.
### Clarifying Questions
- Who calls the control plane: customers through a public API, internal operators, or both?
- What counts as unhealthy for a host: missed heartbeats only, or also resource limits (disk, memory, CPU) and failing local service checks?
- When a host is unhealthy, should the control plane replace it automatically, or only raise an alert?
- How many clusters and hosts must the design handle, and how quickly must a failed host be detected?
- Must running clusters keep working while the control plane itself is down or being upgraded?
### What a Strong Answer Covers
- A durable source of truth for each cluster's desired state, behind an API that records intent and returns before the work finishes
- Reconciliation that drives actual state towards desired state through idempotent, resumable steps, with retries and protection against conflicting operations
- A health-monitoring path, separate from the management path, that scales with the number of hosts
- A failure-detection policy that limits false positives and prevents mass replacement during partitions or monitoring outages
- Availability of the control plane itself, what its outage means for running clusters, and observability
### Follow-up Questions
- A network partition makes a third of the hosts in one zone stop heartbeating at once. What does your system do, and what stops it from replacing all of them?
- How would you roll out a software upgrade across every host of a cluster safely?
- Two requests to resize the same cluster arrive at the same moment. How are they resolved?
- How does the design change when the control plane must manage clusters in several regions?
Overview: Design a control plane that creates, resizes and deletes clusters of cloud hosts, drives each cluster to its desired configuration, and continuously monitors host health. Tests durable desired-state storage, idempotent long-running operations, scalable heartbeat-based failure detection, and safe automatic remediation.
Read the full Snowflake Software Engineer interview experience this question came from