Monitor 100,000 GPU Hosts and Trigger Repair Workflows
Company: Oracle
Role: Site Reliability Engineer
Category: System Design
Difficulty: easy
Interview Round: Onsite
Overview: Design GPU-specific health monitoring and automated repair for a 100,000-host fleet. The solution covers active probes, telemetry ingestion, device-to-fleet aggregation, dashboards, guarded repair state machines, idempotent recovery, blast-radius controls, and failed-workflow escalation.
Read the full Oracle Site Reliability Engineer interview experience this question came from