Cloud Provider Design for Provisioning Resources Through Terraform
Company: Apple
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
You work at a cloud provider, on a team that builds and operates infrastructure. Customers want to provision one of your resources with Terraform: declare it in configuration, run `terraform apply`, and later change or destroy it the same way. From the provider's point of view, what do you need to build so that this works?
Walk through it at a high level: the operations the provider must expose, how Terraform uses them over a resource's lifetime, and how Terraform learns whether a request has actually finished. The question takes about twenty minutes at the end of a mostly conversational screen, so breadth and clear reasoning matter more than depth in any one component.
```hint Walk the lifecycle
Trace one resource through Terraform: the first apply, a later plan with no configuration change, a configuration change, and a destroy. Each step needs something from your side.
```
```hint Not everything is instant
Creating infrastructure can take minutes. Think about what your API should return right away, and how the caller finds out what happened afterwards.
```
### Clarifying Questions
- What kind of resource is being provisioned, and roughly how long does creating or changing one take?
- Does the provider already have a public API, or is designing that API part of the question? Is the Terraform provider plugin itself in scope?
- Which attributes of the resource can be changed in place, and which require replacing it?
- Are there per-customer quotas or rate limits the design must respect?
- Must the design handle resources that were created outside Terraform, for example in a web console?
### What a Strong Answer Covers
- A mapping from Terraform's lifecycle steps (create, refresh, update, destroy, and optionally import) to provider operations.
- How the caller learns that a long-running request has reached a final state, including failures and timeouts.
- Idempotency and retries, so that a repeated or interrupted apply does not create duplicates.
- Read behavior that lets Terraform detect drift and resources deleted outside Terraform.
- Error semantics that separate retryable from permanent failures and give actionable messages.
- The provider-side components behind the API (control plane, state store, workers) at a high level.
### Follow-up Questions
- One `terraform apply` creates hundreds of your resources in parallel. What happens to your API, and what should the provider and the plugin do about rate limits?
- Creation fails halfway: the resource exists but is unhealthy. What should the provider report, and what should Terraform do on the next apply?
- Instead of the client repeatedly asking for status, could the provider push a completion event? What would that change, given that Terraform runs as a short-lived command-line process?
- How would a customer bring an existing resource, created in the console, under Terraform management?
Overview: A cloud infrastructure design question asking what a cloud provider must build so customers can create, update and destroy its resources with Terraform. It tests mapping the Terraform resource lifecycle to provider APIs, handling long-running operations and their status, idempotent retries, and drift detection.
Read the full Apple Software Engineer interview experience this question came from