Design a Strongly Consistent, Fault-Tolerant Database for Fleet Validation Data

Read the full interview experience this question came from →

Quick Overview

System design question about the database between an OS validation service that writes results and a fleet management service that reads them, with strong consistency required. It probes primary and replica failures, synchronous versus asynchronous replication, failover, two-phase commit and consensus.

Design a Strongly Consistent, Fault-Tolerant Database for Fleet Validation Data

Company: Coreweave

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design the database layer between two internal services that manage a large fleet of machines. An OS validation service writes its validation results to the database, and a fleet management service reads them to manage the fleet. Beyond this writer and reader relationship there are almost no functional requirements, and the scale is the whole fleet. The system must be strongly consistent, and the interviewer probed fault tolerance in detail: what happens when the PostgreSQL write primary goes down, how asynchronous or synchronous replication behaves, what happens when a read replica goes down, and where two-phase commit (2PC) and quorum or consensus protocols fit. ### Constraints and Clarifications - One writer service and one reader service; fleet-level scale with no numbers given. State your own estimates. - Strong consistency is a hard requirement. - The discussion assumed a PostgreSQL deployment with a write primary and replicas. - You are not asked to compare database products by workload; focus on consistency and fault tolerance. ### Clarifying Questions - What does a validation record contain (host, OS image or version, outcome, timestamps), and how many hosts and validations per hour does "fleet level" mean? - Does strong consistency mean that every read by the fleet management service reflects every write already acknowledged to the validation service, or is a weaker guarantee acceptable for some reads? - Is losing an acknowledged write ever acceptable? What recovery time is acceptable when the primary fails? - While the primary is unavailable, should the validation service fail, block, or queue its writes? - Is the deployment in one region with several availability zones, or across regions? ### Part 1 — Data model and the consistent read/write path Design the schema and the write and read paths so that the fleet management service always sees a consistent view of validation results. ```hint Not every read is the same Decide which reads drive decisions that must not use stale data and which can tolerate lag, then route each kind accordingly. ``` #### What This Part Should Cover - A schema for validation history and current per-host status, with keys and indexes - Idempotent, retry-safe writes that cannot move a host's status backward - Which reads go to the primary and which to replicas, and how staleness is prevented or bounded - Transaction boundaries and isolation level ### Part 2 — The write primary goes down The PostgreSQL write primary fails. Walk through what happens, and explain how asynchronous versus synchronous replication determines whether acknowledged writes survive and how failover proceeds. ```hint Follow one acknowledged write Take a write that the validation service saw succeed just before the crash, and ask where copies of it existed at that moment. ``` ```hint The old primary comes back Ask what stops the old primary from accepting writes if it returns after a replica has been promoted. ``` #### What This Part Should Cover - Asynchronous versus synchronous replication: durability of acknowledged writes and the latency cost - Failure detection, promotion of the right replica, and fencing to prevent two primaries - Client behavior during failover: retries, idempotency, ambiguous commits and connection routing - The resulting data-loss and recovery-time guarantees ### Part 3 — Replica failures, 2PC and consensus A read replica goes down. What changes for readers, and what changes for the primary? Then explain where two-phase commit and quorum- or consensus-based replication fit, and when you would move beyond a single-primary design. ```hint A synchronous standby is a dependency If the primary waits for a replica before acknowledging, ask what happens to every write when that replica dies. ``` ```hint Two different problems Separate committing one transaction atomically across several databases from keeping several copies of one database in agreement. ``` #### What This Part Should Cover - Loss of an asynchronous replica versus a synchronous standby, and how commits keep flowing - Read routing, lag-based removal of replicas, and capacity planning - What 2PC solves, its blocking failure mode, and whether this design needs it - Quorum and consensus replication (Raft or Paxos based systems, distributed SQL) and their costs ### What a Strong Answer Covers - A precise definition of the consistency required and a design that delivers it - Explicit durability guarantees for acknowledged writes under each failure - Safe failover without split brain - Correct roles for 2PC and for consensus, not used interchangeably - Operations: replication-lag monitoring, backups and point-in-time recovery, failover drills ### Follow-up Questions - The network separates the primary from its synchronous standby but not from the validation service. What should the primary do with incoming writes? - How would you serve strongly consistent reads at a volume one primary cannot handle? - How would the design change if the fleet spans several regions, each with its own validation service? - If you moved to a consensus-based distributed SQL database, what would you gain and what would you pay?

Overview: System design question about the database between an OS validation service that writes results and a fleet management service that reads them, with strong consistency required. It probes primary and replica failures, synchronous versus asynchronous replication, failover, two-phase commit and consensus.

Read the full Coreweave Software Engineer interview experience this question came from

|Home/System Design/Coreweave
Coreweave logo
Coreweave
Sep 29, 2026
mediumSoftware EngineerOnsiteSystem Design
0
0

Design the database layer between two internal services that manage a large fleet of machines. An OS validation service writes its validation results to the database, and a fleet management service reads them to manage the fleet. Beyond this writer and reader relationship there are almost no functional requirements, and the scale is the whole fleet.

The system must be strongly consistent, and the interviewer probed fault tolerance in detail: what happens when the PostgreSQL write primary goes down, how asynchronous or synchronous replication behaves, what happens when a read replica goes down, and where two-phase commit (2PC) and quorum or consensus protocols fit.

Constraints and Clarifications

  • One writer service and one reader service; fleet-level scale with no numbers given. State your own estimates.
  • Strong consistency is a hard requirement.
  • The discussion assumed a PostgreSQL deployment with a write primary and replicas.
  • You are not asked to compare database products by workload; focus on consistency and fault tolerance.

Clarifying Questions Guidance

  • What does a validation record contain (host, OS image or version, outcome, timestamps), and how many hosts and validations per hour does "fleet level" mean?
  • Does strong consistency mean that every read by the fleet management service reflects every write already acknowledged to the validation service, or is a weaker guarantee acceptable for some reads?
  • Is losing an acknowledged write ever acceptable? What recovery time is acceptable when the primary fails?
  • While the primary is unavailable, should the validation service fail, block, or queue its writes?
  • Is the deployment in one region with several availability zones, or across regions?

Part 1 — Data model and the consistent read/write path

Design the schema and the write and read paths so that the fleet management service always sees a consistent view of validation results.

What This Part Should Cover Guidance

  • A schema for validation history and current per-host status, with keys and indexes
  • Idempotent, retry-safe writes that cannot move a host's status backward
  • Which reads go to the primary and which to replicas, and how staleness is prevented or bounded
  • Transaction boundaries and isolation level

Part 2 — The write primary goes down

The PostgreSQL write primary fails. Walk through what happens, and explain how asynchronous versus synchronous replication determines whether acknowledged writes survive and how failover proceeds.

What This Part Should Cover Guidance

  • Asynchronous versus synchronous replication: durability of acknowledged writes and the latency cost
  • Failure detection, promotion of the right replica, and fencing to prevent two primaries
  • Client behavior during failover: retries, idempotency, ambiguous commits and connection routing
  • The resulting data-loss and recovery-time guarantees

Part 3 — Replica failures, 2PC and consensus

A read replica goes down. What changes for readers, and what changes for the primary? Then explain where two-phase commit and quorum- or consensus-based replication fit, and when you would move beyond a single-primary design.

What This Part Should Cover Guidance

  • Loss of an asynchronous replica versus a synchronous standby, and how commits keep flowing
  • Read routing, lag-based removal of replicas, and capacity planning
  • What 2PC solves, its blocking failure mode, and whether this design needs it
  • Quorum and consensus replication (Raft or Paxos based systems, distributed SQL) and their costs

What a Strong Answer Covers Guidance

  • A precise definition of the consistency required and a design that delivers it
  • Explicit durability guarantees for acknowledged writes under each failure
  • Safe failover without split brain
  • Correct roles for 2PC and for consensus, not used interchangeably
  • Operations: replication-lag monitoring, backups and point-in-time recovery, failover drills

Follow-up Questions Guidance

  • The network separates the primary from its synchronous standby but not from the validation service. What should the primary do with incoming writes?
  • How would you serve strongly consistent reads at a volume one primary cannot handle?
  • How would the design change if the fleet spans several regions, each with its own validation service?
  • If you moved to a consensus-based distributed SQL database, what would you gain and what would you pay?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...