Design a metrics monitoring system, going deep on one or two components

Read the full interview experience this question came from →

Quick Overview

A system design question about a metrics monitoring system that collects time series from services and hosts, stores them, serves dashboard queries and raises alerts. After a brief end-to-end sketch, the candidate goes deep on one or two areas, such as time-series storage, label cardinality, downsampling or reliable alert evaluation.

Design a metrics monitoring system, going deep on one or two components

Company: OpenAI

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design a metrics monitoring system. Services and hosts across a fleet emit numeric measurements over time, such as request counts, error counts, latencies, and CPU and memory use. The system collects these measurements, stores them as time series, lets engineers query and chart them on dashboards, and alerts the on-call engineer when a rule is breached. You are not expected to design every component in full. Sketch the end-to-end pipeline briefly, then choose one or two areas and design them in depth. ```hint Start from one data point Write down exactly what a single measurement looks like, including the labels that identify where it came from. Many of the hardest scaling problems come from how many distinct label combinations exist, not from how many points arrive. ``` ```hint Choose the deep dive deliberately Before going deep, say which area holds the hardest problem at the scale you estimated, and why. A deep dive the interviewer can follow beats a shallow tour of every box. ``` ```hint Old data is read differently Recent data is read at full detail by dashboards and alerts; data from months ago is mostly viewed over long ranges. Consider what that difference allows you to do. ``` ### Clarifying Questions - Is there a specific twist on the usual monitoring problem, and which part of the pipeline should the deep dive cover: collection, storage, querying, or alerting? - How many hosts and services are monitored, how many metrics does each emit, and how often? Roughly how many distinct time series are active at once? - Do services push measurements to the system, or does the system pull them from each service? - How long must data be kept, and at what resolution once it is old? - How fresh must dashboards and alerts be, and is losing a small fraction of data points acceptable? - Are logs and traces in scope, or only numeric metrics? - How should alerts reach people, and are grouping, silencing and escalation required? ### What a Strong Answer Covers - A precise data model for a measurement and a series, and estimates of ingest rate, active series count and storage - A brief end-to-end pipeline, followed by a justified choice of deep-dive areas - Storage suited to time-series access: append-heavy writes, compression, time partitioning, retention and downsampling - Handling of label cardinality and ingest bursts, including what is rejected or dropped under overload - Alerting that stays correct and available when data is missing or parts of the system fail - How the monitoring system itself is monitored, and the trade-offs behind push or pull and freshness versus cost ### Follow-up Questions - A team adds a label holding a user ID, and the number of series jumps by orders of magnitude. How does your system detect and contain it? - How do you compute a fleet-wide 99th-percentile latency when each host reports its own measurements? - Alerts must keep firing even when the main storage cluster is unavailable. What changes? - How would you keep a year of data cheaply while keeping queries over the last hour fast?

Overview: A system design question about a metrics monitoring system that collects time series from services and hosts, stores them, serves dashboard queries and raises alerts. After a brief end-to-end sketch, the candidate goes deep on one or two areas, such as time-series storage, label cardinality, downsampling or reliable alert evaluation.

Read the full OpenAI Software Engineer interview experience this question came from

|Home/System Design/OpenAI
OpenAI logo
OpenAI
Oct 1, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Design a metrics monitoring system. Services and hosts across a fleet emit numeric measurements over time, such as request counts, error counts, latencies, and CPU and memory use. The system collects these measurements, stores them as time series, lets engineers query and chart them on dashboards, and alerts the on-call engineer when a rule is breached.

You are not expected to design every component in full. Sketch the end-to-end pipeline briefly, then choose one or two areas and design them in depth.

Clarifying Questions Guidance

  • Is there a specific twist on the usual monitoring problem, and which part of the pipeline should the deep dive cover: collection, storage, querying, or alerting?
  • How many hosts and services are monitored, how many metrics does each emit, and how often? Roughly how many distinct time series are active at once?
  • Do services push measurements to the system, or does the system pull them from each service?
  • How long must data be kept, and at what resolution once it is old?
  • How fresh must dashboards and alerts be, and is losing a small fraction of data points acceptable?
  • Are logs and traces in scope, or only numeric metrics?
  • How should alerts reach people, and are grouping, silencing and escalation required?

What a Strong Answer Covers Guidance

  • A precise data model for a measurement and a series, and estimates of ingest rate, active series count and storage
  • A brief end-to-end pipeline, followed by a justified choice of deep-dive areas
  • Storage suited to time-series access: append-heavy writes, compression, time partitioning, retention and downsampling
  • Handling of label cardinality and ingest bursts, including what is rejected or dropped under overload
  • Alerting that stays correct and available when data is missing or parts of the system fail
  • How the monitoring system itself is monitored, and the trade-offs behind push or pull and freshness versus cost

Follow-up Questions Guidance

  • A team adds a label holding a user ID, and the number of series jumps by orders of magnitude. How does your system detect and contain it?
  • How do you compute a fleet-wide 99th-percentile latency when each host reports its own measurements?
  • Alerts must keep firing even when the main storage cluster is unavailable. What changes?
  • How would you keep a year of data cheaply while keeping queries over the last hour fast?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...