Design a Lab Machine Fleet Monitor: Payload Dispatch, Hang Detection, and Remote Reboot

Read the full interview experience this question came from →

Quick Overview

Design a system that dispatches compute payloads to lab machines that poll for work and shows on a central dashboard whether each payload is downloading, running, completed, crashed, or hung. It tests agent design, hang detection, crash and hang logs, authorized reboot and power cycle, reruns after interruption, and refreshable machine inventory.

Design a Lab Machine Fleet Monitor: Payload Dispatch, Hang Detection, and Remote Reboot

Company: Microsoft

Role: Software Engineer

Category: System Design

Difficulty: easy

Interview Round: Onsite

Design a system that lets users monitor a collection of lab machines that run compute loads. - Machines periodically ask the system for work. A work item identifies a compute payload that the machine must download, verify, start, and then monitor. - A central dashboard shows whether each payload is downloading, running, completed, crashed, or appears hung. For crashed and hung payloads, operators need useful logs. - An authorized operator may request a reboot or a power cycle of a machine. - A payload that is interrupted starts again from the beginning after the machine recovers. - Payload results must be available from the portal. - Users can see each machine's OS, CPU, and state. This information updates every 2 minutes, and a user can also force a refresh. No production code or cloud-specific service names are required. Draw the main components as you discuss them. ```hint Who can reach whom Machines start every conversation by asking for work. Consider how a reboot request or a forced refresh reaches a machine between two of its check-ins, and what still works when the machine's own software is frozen. ``` ```hint Hung is not the same as silent If the dashboard only notices that reports stopped, it cannot tell a hung payload from a crashed agent, a dead machine, or a network problem. Decide which component can observe each of those conditions. ``` ```hint Runs, not just work items Because an interrupted payload starts over, one work item can run more than once. Think about how the status, logs, and results of an interrupted run stay separate from the run that replaces it. ``` ### Constraints and Clarifications - Inventory (OS, CPU, machine state) refreshes on a 2-minute cycle, plus on demand. - An interrupted payload is rerun from the start. It does not resume from a checkpoint. - Only authorized operators may reboot or power cycle a machine. - Assume you may install software of your own design on every lab machine. - No scale numbers were given. State any you assume. ### Clarifying Questions - How many machines are there, how many payloads run on one machine at once, and how large are payloads, logs, and results? - Does a payload emit any progress or heartbeat signal, or must "appears hung" be inferred from outside the process? Is the hang threshold the same for every payload? - Can the system open connections to the machines, or can only the machines connect outward (for example, from behind a lab firewall)? - Do the machines have out-of-band management hardware, such as a baseboard management controller or a switched power unit, that can power cycle them when the operating system is unresponsive? - How quickly must a forced refresh show new data, and how long must logs and results be retained? - Who creates work items, and can a work item target a specific machine or any machine with a suitable OS and CPU? ### What a Strong Answer Covers - A component diagram: machine agent, work dispatch, status ingestion, payload and result storage, log storage, command path, and the dashboard and portal - A payload lifecycle with explicit transitions for download, verification, start, completion, crash, hang, and interruption, including how a rerun is represented - A precise definition of "appears hung", and how it is distinguished from a machine or agent that is offline - A command path for reboot, power cycle, and forced refresh that works between check-ins and when the machine is unresponsive, with authorization and auditing - Log capture for crashes and hangs that survives the reboot that often follows - Freshness and load reasoning for the 2-minute inventory updates and on-demand refreshes - Failure handling: lost reports, duplicate or replayed messages, agent restarts, and network partitions ### Follow-up Questions - An operator power cycles a machine while its agent is uploading a result. What does the dashboard show, and what keeps a partial result out of the portal? - How would you roll out a new version of the machine agent across the fleet without disrupting running payloads? - How do you make sure a downloaded payload is the one the work item intended, and what happens if verification keeps failing on one machine? - Hundreds of users press force refresh on the same machine at once. How do you protect the machine and the backend?

Overview: Design a system that dispatches compute payloads to lab machines that poll for work and shows on a central dashboard whether each payload is downloading, running, completed, crashed, or hung. It tests agent design, hang detection, crash and hang logs, authorized reboot and power cycle, reruns after interruption, and refreshable machine inventory.

Read the full Microsoft Software Engineer interview experience this question came from

|Home/System Design/Microsoft
Microsoft logo
Microsoft
Sep 30, 2026
easySoftware EngineerOnsiteSystem Design
0
0

Design a system that lets users monitor a collection of lab machines that run compute loads.

  • Machines periodically ask the system for work. A work item identifies a compute payload that the machine must download, verify, start, and then monitor.
  • A central dashboard shows whether each payload is downloading, running, completed, crashed, or appears hung. For crashed and hung payloads, operators need useful logs.
  • An authorized operator may request a reboot or a power cycle of a machine.
  • A payload that is interrupted starts again from the beginning after the machine recovers.
  • Payload results must be available from the portal.
  • Users can see each machine's OS, CPU, and state. This information updates every 2 minutes, and a user can also force a refresh.

No production code or cloud-specific service names are required. Draw the main components as you discuss them.

Constraints and Clarifications

  • Inventory (OS, CPU, machine state) refreshes on a 2-minute cycle, plus on demand.
  • An interrupted payload is rerun from the start. It does not resume from a checkpoint.
  • Only authorized operators may reboot or power cycle a machine.
  • Assume you may install software of your own design on every lab machine.
  • No scale numbers were given. State any you assume.

Clarifying Questions Guidance

  • How many machines are there, how many payloads run on one machine at once, and how large are payloads, logs, and results?
  • Does a payload emit any progress or heartbeat signal, or must "appears hung" be inferred from outside the process? Is the hang threshold the same for every payload?
  • Can the system open connections to the machines, or can only the machines connect outward (for example, from behind a lab firewall)?
  • Do the machines have out-of-band management hardware, such as a baseboard management controller or a switched power unit, that can power cycle them when the operating system is unresponsive?
  • How quickly must a forced refresh show new data, and how long must logs and results be retained?
  • Who creates work items, and can a work item target a specific machine or any machine with a suitable OS and CPU?

What a Strong Answer Covers Guidance

  • A component diagram: machine agent, work dispatch, status ingestion, payload and result storage, log storage, command path, and the dashboard and portal
  • A payload lifecycle with explicit transitions for download, verification, start, completion, crash, hang, and interruption, including how a rerun is represented
  • A precise definition of "appears hung", and how it is distinguished from a machine or agent that is offline
  • A command path for reboot, power cycle, and forced refresh that works between check-ins and when the machine is unresponsive, with authorization and auditing
  • Log capture for crashes and hangs that survives the reboot that often follows
  • Freshness and load reasoning for the 2-minute inventory updates and on-demand refreshes
  • Failure handling: lost reports, duplicate or replayed messages, agent restarts, and network partitions

Follow-up Questions Guidance

  • An operator power cycles a machine while its agent is uploading a result. What does the dashboard show, and what keeps a partial result out of the portal?
  • How would you roll out a new version of the machine agent across the fleet without disrupting running payloads?
  • How do you make sure a downloaded payload is the one the work item intended, and what happens if verification keeps failing on one machine?
  • Hundreds of users press force refresh on the same machine at once. How do you protect the machine and the backend?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...