Design power grid device monitoring over an unreliable public internet

Quick Overview

Design a monitoring system for power grid field devices that report their status over an unreliable public internet. It tests store-and-forward reporting, idempotent ingestion of late and duplicate data, telling silent devices from failed ones, alerting, and security for critical infrastructure.

Design power grid device monitoring over an unreliable public internet

Company: OpenAI

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a system that monitors devices on a power grid. Devices in the field report their status to a central monitoring system over the public internet, and the public internet is unreliable: connections drop, messages are delayed or repeated, and a device can stay unreachable for long periods. Operators use the system to see the current state of the grid and to be alerted when something goes wrong. ```hint Silence is ambiguous When a device stops reporting, list everything that could explain it, and what the system would need to tell those cases apart. ``` ```hint Where data can be lost Trace one status report from the device to the operator's screen, and mark every point where it can be dropped, duplicated or delayed. ``` ### Clarifying Questions - How many devices are there, how often does each one report, and what does a status report contain? - Is the system for monitoring only, or must it also send commands back to devices? - How quickly must an operator learn that a device has failed or gone silent? - How much storage and compute do the devices have for buffering? - How long must status history be kept, and at what resolution? ### What a Strong Answer Covers - Device-side behavior that survives disconnection: local buffering, sequence numbers, and retries with backoff and jitter - Idempotent, ordered ingestion that tolerates duplicates, late data and reconnect storms - A clear distinction between a device reporting a fault, a device that is unreachable, and a regional network outage - Separate stores for current state and history, sized from estimates - Alerting that avoids floods and false alarms, and monitoring of the monitoring system itself - Security suited to critical infrastructure exposed on the public internet ### Follow-up Questions - A regional internet outage takes many devices offline for an hour, and then they all reconnect at once. Walk through what your system does, step by step. - If operators must also send commands to devices, what changes in the protocol, the architecture and the safety requirements? - How would you detect a device that keeps reporting on schedule but whose readings are wrong?

Overview: Design a monitoring system for power grid field devices that report their status over an unreliable public internet. It tests store-and-forward reporting, idempotent ingestion of late and duplicate data, telling silent devices from failed ones, alerting, and security for critical infrastructure.

|Home/System Design/OpenAI
OpenAI logo
OpenAI
Sep 28, 2026
mediumSoftware EngineerOnsiteSystem Design
1
0

Design a system that monitors devices on a power grid. Devices in the field report their status to a central monitoring system over the public internet, and the public internet is unreliable: connections drop, messages are delayed or repeated, and a device can stay unreachable for long periods. Operators use the system to see the current state of the grid and to be alerted when something goes wrong.

Clarifying Questions Guidance

  • How many devices are there, how often does each one report, and what does a status report contain?
  • Is the system for monitoring only, or must it also send commands back to devices?
  • How quickly must an operator learn that a device has failed or gone silent?
  • How much storage and compute do the devices have for buffering?
  • How long must status history be kept, and at what resolution?

What a Strong Answer Covers Guidance

  • Device-side behavior that survives disconnection: local buffering, sequence numbers, and retries with backoff and jitter
  • Idempotent, ordered ingestion that tolerates duplicates, late data and reconnect storms
  • A clear distinction between a device reporting a fault, a device that is unreachable, and a regional network outage
  • Separate stores for current state and history, sized from estimates
  • Alerting that avoids floods and false alarms, and monitoring of the monitoring system itself
  • Security suited to critical infrastructure exposed on the public internet

Follow-up Questions Guidance

  • A regional internet outage takes many devices offline for an hour, and then they all reconnect at once. Walk through what your system does, step by step.
  • If operators must also send commands to devices, what changes in the protocol, the architecture and the safety requirements?
  • How would you detect a device that keeps reporting on schedule but whose readings are wrong?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...