Design a Devbox Service for On-Demand Remote Development Environments

Read the full interview experience this question came from →

Quick Overview

System design question asking you to design a devbox service that gives engineers on-demand remote development environments. It tests scoping requirements before architecture, the devbox lifecycle, fast provisioning, persistent workspaces, idle cost control, isolation, connectivity and failure recovery.

Design a Devbox Service for On-Demand Remote Development Environments

Company: OpenAI

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Technical Screen

Design a devbox service: a system that lets engineers get a remote development environment (a "devbox") on demand, work in it from their own laptop, and have it cleaned up when it is no longer needed. The prompt is deliberately short, and the interviewer gives little direction beyond it. Working out what exactly needs to be built is part of the exercise: a design that turns into a general-purpose container orchestration platform, without the devbox lifecycle an engineer actually experiences, misses the question. ```hint Start from one engineer's day Before drawing a scheduler, pin down what a single engineer experiences: how a devbox is requested, how long until it is usable, how they connect, what survives a stop or a crash, and when it goes away. Let the infrastructure follow from those answers. ``` ```hint Disposable versus personal Decide which parts of a devbox can be thrown away and rebuilt, and which belong to the engineer and must persist. That split drives provisioning speed, cost and recovery. ``` ### Constraints and Clarifications - Format: a system design phone screen of just under an hour, of which about 10 minutes are reserved for the candidate's questions. - No scale, latency target or technology is given; agree on them with the interviewer. ### Clarifying Questions - Who are the users, and roughly how many devboxes exist and run at the same time? - Is a devbox a container or a full virtual machine? Do some users need large machines or GPUs? - How do engineers connect: SSH, a local editor with a remote-development extension, or a browser IDE? - What must persist across a stop and restart: the home directory, uncommitted changes in the repository, installed packages, running processes? - How long may it take from request to a usable devbox, for a new devbox and for resuming a stopped one? - Should idle devboxes stop automatically, and are there per-user quotas or cost limits? - What may a devbox access: source repositories, package mirrors, internal services, secrets, production data? ### What a Strong Answer Covers - Scoped functional requirements for the devbox lifecycle (create, connect, stop, resume, delete, idle shutdown), agreed with the interviewer before any architecture. - A control plane with an API, a per-devbox state machine, placement onto a compute pool and reconciliation of desired versus actual state; and how a user's connection is routed to the right machine. - Fast provisioning for large repositories: prebuilt images, cached layers, warm capacity, snapshot-based workspaces. - Persistence and lifecycle: compute separated from the engineer's persistent workspace, idle detection, quotas and cost control. - Isolation and security: per-user isolation, authenticated access through a gateway, short-lived credentials, network egress policy. - Failure handling and observability: host failure, stuck provisioning, control plane outage, and the metrics that show engineers are productive. ### Follow-up Questions - A host running many devboxes dies. What does each affected engineer lose, and how quickly can they be working again? - How would you reduce the time to a usable devbox from minutes to seconds for a very large monorepo? - How would you support devboxes that need GPUs when that capacity is scarce and expensive? - How would you guarantee that a devbox can reach source control and package mirrors but not production data?

Overview: System design question asking you to design a devbox service that gives engineers on-demand remote development environments. It tests scoping requirements before architecture, the devbox lifecycle, fast provisioning, persistent workspaces, idle cost control, isolation, connectivity and failure recovery.

Read the full OpenAI Software Engineer interview experience this question came from

|Home/System Design/OpenAI
OpenAI logo
OpenAI
Sep 8, 2026
hardSoftware EngineerTechnical ScreenSystem Design
0
0

Design a devbox service: a system that lets engineers get a remote development environment (a "devbox") on demand, work in it from their own laptop, and have it cleaned up when it is no longer needed.

The prompt is deliberately short, and the interviewer gives little direction beyond it. Working out what exactly needs to be built is part of the exercise: a design that turns into a general-purpose container orchestration platform, without the devbox lifecycle an engineer actually experiences, misses the question.

Constraints and Clarifications

  • Format: a system design phone screen of just under an hour, of which about 10 minutes are reserved for the candidate's questions.
  • No scale, latency target or technology is given; agree on them with the interviewer.

Clarifying Questions Guidance

  • Who are the users, and roughly how many devboxes exist and run at the same time?
  • Is a devbox a container or a full virtual machine? Do some users need large machines or GPUs?
  • How do engineers connect: SSH, a local editor with a remote-development extension, or a browser IDE?
  • What must persist across a stop and restart: the home directory, uncommitted changes in the repository, installed packages, running processes?
  • How long may it take from request to a usable devbox, for a new devbox and for resuming a stopped one?
  • Should idle devboxes stop automatically, and are there per-user quotas or cost limits?
  • What may a devbox access: source repositories, package mirrors, internal services, secrets, production data?

What a Strong Answer Covers Guidance

  • Scoped functional requirements for the devbox lifecycle (create, connect, stop, resume, delete, idle shutdown), agreed with the interviewer before any architecture.
  • A control plane with an API, a per-devbox state machine, placement onto a compute pool and reconciliation of desired versus actual state; and how a user's connection is routed to the right machine.
  • Fast provisioning for large repositories: prebuilt images, cached layers, warm capacity, snapshot-based workspaces.
  • Persistence and lifecycle: compute separated from the engineer's persistent workspace, idle detection, quotas and cost control.
  • Isolation and security: per-user isolation, authenticated access through a gateway, short-lived credentials, network egress policy.
  • Failure handling and observability: host failure, stuck provisioning, control plane outage, and the metrics that show engineers are productive.

Follow-up Questions Guidance

  • A host running many devboxes dies. What does each affected engineer lose, and how quickly can they be working again?
  • How would you reduce the time to a usable devbox from minutes to seconds for a very large monorepo?
  • How would you support devboxes that need GPUs when that capacity is scarce and expensive?
  • How would you guarantee that a devbox can reach source control and package mirrors but not production data?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...