Thread-Safe Durable Event Writer Shared by Thousands of Threads
Quick Overview
Design a thread-safe DataWriter whose push method durably appends opaque records of up to 1 KB to one local file while thousands of threads call it at once. Tests durability contracts, batching disk syncs across callers, crash-safe record framing and handling of failed writes.
Thread-Safe Durable Event Writer Shared by Thousands of Threads
Company: Databricks
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a thread-safe, durable event writer used by a single application running on one server. Thousands of application threads share a single instance of this class and continuously call it to append data to a file on local disk.
```python
class DataWriter:
def __init__(self, file_path_on_disk: str): ...
def push(self, data: bytes) -> None: ...
```
Each `data` value is at most 1 KB. Its contents are opaque: you do not need to interpret or process them. The interviewer opens with: "Go ahead and start designing. Ask me anything you need before you begin." A closely related version of this question states the core tension directly: writes to the single file must be synchronous, yet many threads write at the same time.
```hint What does "durable" promise?
Pin down exactly what must be true when `push` returns. Then count how many expensive disk operations that promise costs per call in a naive design.
```
```hint Many callers, one disk
Look for a way for one expensive operation to satisfy many waiting callers at once.
```
### Clarifying Questions
- When `push` returns, must the data be on stable storage (able to survive power loss), or is handing it to the operating system enough?
- Must records from different threads appear in the file in a particular order, and must each thread's own records stay in order?
- Will something read this file back later, and must it be able to split the file into the original records?
- What throughput and per-call latency are expected, and may `push` block under load?
- What should happen after a disk error: fail every later call, or keep trying?
- Is the file ever rotated or truncated, and does the class need a `close`?
### What a Strong Answer Covers
- A precise durability contract for `push`, and how the design meets it
- A correct concurrency design (no lost, duplicated or interleaved records) that avoids one disk sync per call
- An on-disk record format that lets a reader recover record boundaries and detect a torn or corrupt tail after a crash
- Latency versus throughput analysis of batching, and backpressure when the disk falls behind
- Error handling for failed writes or syncs, and clean shutdown
- Working code or precise pseudocode for `push` and the flushing logic
### Follow-up Questions
- How would the design change if `push` could return before the data is durable, with a separate way to wait for durability?
- How would you rotate the file without blocking writers for long?
- How would you test that no acknowledged record is ever lost after a crash?
Overview: Design a thread-safe DataWriter whose push method durably appends opaque records of up to 1 KB to one local file while thousands of threads call it at once. Tests durability contracts, batching disk syncs across callers, crash-safe record framing and handling of failed writes.
Thread-Safe Durable Event Writer Shared by Thousands of Threads
Databricks
Sep 11, 2026
mediumSoftware EngineerOnsiteSystem Design
1
0
Design a thread-safe, durable event writer used by a single application running on one server. Thousands of application threads share a single instance of this class and continuously call it to append data to a file on local disk.
Each data value is at most 1 KB. Its contents are opaque: you do not need to interpret or process them. The interviewer opens with: "Go ahead and start designing. Ask me anything you need before you begin." A closely related version of this question states the core tension directly: writes to the single file must be synchronous, yet many threads write at the same time.
Clarifying Questions Guidance
When
push
returns, must the data be on stable storage (able to survive power loss), or is handing it to the operating system enough?
Must records from different threads appear in the file in a particular order, and must each thread's own records stay in order?
Will something read this file back later, and must it be able to split the file into the original records?
What throughput and per-call latency are expected, and may
push
block under load?
What should happen after a disk error: fail every later call, or keep trying?
Is the file ever rotated or truncated, and does the class need a
close
?
What a Strong Answer Covers Guidance
A precise durability contract for
push
, and how the design meets it
A correct concurrency design (no lost, duplicated or interleaved records) that avoids one disk sync per call
An on-disk record format that lets a reader recover record boundaries and detect a torn or corrupt tail after a crash
Latency versus throughput analysis of batching, and backpressure when the disk falls behind
Error handling for failed writes or syncs, and clean shutdown
Working code or precise pseudocode for
push
and the flushing logic
Follow-up Questions Guidance
How would the design change if
push
could return before the data is durable, with a separate way to wait for durability?
How would you rotate the file without blocking writers for long?
How would you test that no acknowledged record is ever lost after a crash?