Monitor 100,000 GPU Hosts and Trigger Repair Workflows

Read the full interview experience this question came from →

Quick Overview

Design GPU-specific health monitoring and automated repair for a 100,000-host fleet. The solution covers active probes, telemetry ingestion, device-to-fleet aggregation, dashboards, guarded repair state machines, idempotent recovery, blast-radius controls, and failed-workflow escalation.

Monitor 100,000 GPU Hosts and Trigger Repair Workflows

Company: Oracle

Role: Site Reliability Engineer

Category: System Design

Difficulty: easy

Interview Round: Onsite

Overview: Design GPU-specific health monitoring and automated repair for a 100,000-host fleet. The solution covers active probes, telemetry ingestion, device-to-fleet aggregation, dashboards, guarded repair state machines, idempotent recovery, blast-radius controls, and failed-workflow escalation.

Read the full Oracle Site Reliability Engineer interview experience this question came from

|Home/System Design/Oracle
Oracle logo
Oracle
May 24, 2026
easySite Reliability EngineerOnsiteSystem Design
0
0
Loading...

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...