Thread-Safe Durable Event Writer Shared by Thousands of Threads

Quick Overview

Design a thread-safe DataWriter whose push method durably appends opaque records of up to 1 KB to one local file while thousands of threads call it at once. Tests durability contracts, batching disk syncs across callers, crash-safe record framing and handling of failed writes.

Thread-Safe Durable Event Writer Shared by Thousands of Threads

Company: Databricks

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a thread-safe, durable event writer used by a single application running on one server. Thousands of application threads share a single instance of this class and continuously call it to append data to a file on local disk. ```python class DataWriter: def __init__(self, file_path_on_disk: str): ... def push(self, data: bytes) -> None: ... ``` Each `data` value is at most 1 KB. Its contents are opaque: you do not need to interpret or process them. The interviewer opens with: "Go ahead and start designing. Ask me anything you need before you begin." A closely related version of this question states the core tension directly: writes to the single file must be synchronous, yet many threads write at the same time. ```hint What does "durable" promise? Pin down exactly what must be true when `push` returns. Then count how many expensive disk operations that promise costs per call in a naive design. ``` ```hint Many callers, one disk Look for a way for one expensive operation to satisfy many waiting callers at once. ``` ### Clarifying Questions - When `push` returns, must the data be on stable storage (able to survive power loss), or is handing it to the operating system enough? - Must records from different threads appear in the file in a particular order, and must each thread's own records stay in order? - Will something read this file back later, and must it be able to split the file into the original records? - What throughput and per-call latency are expected, and may `push` block under load? - What should happen after a disk error: fail every later call, or keep trying? - Is the file ever rotated or truncated, and does the class need a `close`? ### What a Strong Answer Covers - A precise durability contract for `push`, and how the design meets it - A correct concurrency design (no lost, duplicated or interleaved records) that avoids one disk sync per call - An on-disk record format that lets a reader recover record boundaries and detect a torn or corrupt tail after a crash - Latency versus throughput analysis of batching, and backpressure when the disk falls behind - Error handling for failed writes or syncs, and clean shutdown - Working code or precise pseudocode for `push` and the flushing logic ### Follow-up Questions - How would the design change if `push` could return before the data is durable, with a separate way to wait for durability? - How would you rotate the file without blocking writers for long? - How would you test that no acknowledged record is ever lost after a crash?

Overview: Design a thread-safe DataWriter whose push method durably appends opaque records of up to 1 KB to one local file while thousands of threads call it at once. Tests durability contracts, batching disk syncs across callers, crash-safe record framing and handling of failed writes.

|Home/System Design/Databricks
Databricks logo
Databricks
Sep 11, 2026
mediumSoftware EngineerOnsiteSystem Design
1
0

Design a thread-safe, durable event writer used by a single application running on one server. Thousands of application threads share a single instance of this class and continuously call it to append data to a file on local disk.

class DataWriter:
    def __init__(self, file_path_on_disk: str): ...
    def push(self, data: bytes) -> None: ...

Each data value is at most 1 KB. Its contents are opaque: you do not need to interpret or process them. The interviewer opens with: "Go ahead and start designing. Ask me anything you need before you begin." A closely related version of this question states the core tension directly: writes to the single file must be synchronous, yet many threads write at the same time.

Clarifying Questions Guidance

  • When push returns, must the data be on stable storage (able to survive power loss), or is handing it to the operating system enough?
  • Must records from different threads appear in the file in a particular order, and must each thread's own records stay in order?
  • Will something read this file back later, and must it be able to split the file into the original records?
  • What throughput and per-call latency are expected, and may push block under load?
  • What should happen after a disk error: fail every later call, or keep trying?
  • Is the file ever rotated or truncated, and does the class need a close ?

What a Strong Answer Covers Guidance

  • A precise durability contract for push , and how the design meets it
  • A correct concurrency design (no lost, duplicated or interleaved records) that avoids one disk sync per call
  • An on-disk record format that lets a reader recover record boundaries and detect a torn or corrupt tail after a crash
  • Latency versus throughput analysis of batching, and backpressure when the disk falls behind
  • Error handling for failed writes or syncs, and clean shutdown
  • Working code or precise pseudocode for push and the flushing logic

Follow-up Questions Guidance

  • How would the design change if push could return before the data is durable, with a separate way to wait for durability?
  • How would you rotate the file without blocking writers for long?
  • How would you test that no acknowledged record is ever lost after a crash?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...