Design a Distributed File System, Scoping the Requirements Yourself

Quick Overview

Design a distributed file system from a one-line prompt, driving the scoping of workload, semantics, scale and durability yourself. Tests separating metadata from data, block replication and placement, read and write paths, failure recovery, and scaling the namespace for a staff-level backend role.

Design a Distributed File System, Scoping the Requirements Yourself

Company: Databricks

Role: Backend Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a distributed file system. This was a phone-screen system design question for a staff-level backend role. The report gives only the prompt itself, with no scale numbers, requirements or follow-ups. Scoping is therefore part of the exercise: you are expected to drive the requirements discussion before you design. ```hint Two very different kinds of state Consider separating the state that describes files and directories from the bytes inside the files. The two differ in size, access pattern and consistency needs. ``` ```hint Machines will fail Assume disks and servers fail routinely. Decide how the system notices a failure, and how data regains its redundancy without an operator. ``` ### Clarifying Questions - Which workload dominates: a few very large files that are read and appended in bulk, or many small files with random reads and writes? - Which semantics are required: full POSIX (in-place overwrites, rename, locks) or a simpler model such as write-once or append-only files? - What scale should the design target in total bytes, number of files and concurrent clients? - What are the durability and availability targets? Must the system survive the loss of a rack, or of a whole data center? - Is strong consistency needed for reads after writes, and for directory operations such as rename? ### What a Strong Answer Covers - Requirements and scale stated explicitly before the design, with design choices tied back to them - A metadata layer (namespace, file-to-block mapping, block locations), and how it is made durable, highly available and scalable - A data layer: block size, placement, replication or erasure coding, and awareness of failure domains - Read, write and append paths, including who coordinates concurrent writers and what consistency clients observe - Failure detection, re-replication, data integrity checks, and recovery from a failed metadata node - Operational concerns: the many-small-files problem, rebalancing, garbage collection of deleted blocks, and monitoring ### Follow-up Questions - The metadata server becomes the bottleneck at billions of files. How do you scale the namespace? - How would you support cheap snapshots of a directory tree? - When would you choose erasure coding over replication, and what does it cost on reads and repairs? - What happens if a client crashes in the middle of writing a block?

Overview: Design a distributed file system from a one-line prompt, driving the scoping of workload, semantics, scale and durability yourself. Tests separating metadata from data, block replication and placement, read and write paths, failure recovery, and scaling the namespace for a staff-level backend role.

|Home/System Design/Databricks
Databricks logo
Databricks
Sep 11, 2026
mediumBackend EngineerOnsiteSystem Design
1
0

Design a distributed file system.

This was a phone-screen system design question for a staff-level backend role. The report gives only the prompt itself, with no scale numbers, requirements or follow-ups. Scoping is therefore part of the exercise: you are expected to drive the requirements discussion before you design.

Clarifying Questions Guidance

  • Which workload dominates: a few very large files that are read and appended in bulk, or many small files with random reads and writes?
  • Which semantics are required: full POSIX (in-place overwrites, rename, locks) or a simpler model such as write-once or append-only files?
  • What scale should the design target in total bytes, number of files and concurrent clients?
  • What are the durability and availability targets? Must the system survive the loss of a rack, or of a whole data center?
  • Is strong consistency needed for reads after writes, and for directory operations such as rename?

What a Strong Answer Covers Guidance

  • Requirements and scale stated explicitly before the design, with design choices tied back to them
  • A metadata layer (namespace, file-to-block mapping, block locations), and how it is made durable, highly available and scalable
  • A data layer: block size, placement, replication or erasure coding, and awareness of failure domains
  • Read, write and append paths, including who coordinates concurrent writers and what consistency clients observe
  • Failure detection, re-replication, data integrity checks, and recovery from a failed metadata node
  • Operational concerns: the many-small-files problem, rebalancing, garbage collection of deleted blocks, and monitoring

Follow-up Questions Guidance

  • The metadata server becomes the bottleneck at billions of files. How do you scale the namespace?
  • How would you support cheap snapshots of a directory tree?
  • When would you choose erasure coding over replication, and what does it cost on reads and repairs?
  • What happens if a client crashes in the middle of writing a block?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...