Find Duplicate Files by Content with Efficient I/O

Read the full interview experience this question came from →

Quick Overview

Find content-identical files with size filters, streaming digests, exact collision checks, bounded I/O, and explicit handling of changing or unreadable files.

Find Duplicate Files by Content with Efficient I/O

Company: Salient Ai

Role: Software Engineer

Category: Software Engineering Fundamentals

Difficulty: hard

Interview Round: Onsite

Using supplied filesystem helpers, design an operation that finds files with identical contents. Explain how to reduce unnecessary file reads and memory use while preserving the distinction between exact content equality and a likely match based on metadata or hashing. ### Constraints & Assumptions - The source names content-based duplicate-file detection and follow-ups on I/O and reading optimization. It does not supply the exact helper signatures. - **Practice scope:** report groups containing at least two paths whose complete byte contents are equal. Do not delete or overwrite files as part of duplicate detection. - Ask which helpers provide enumeration, size, bounded reads, and stable file identity or version. Use only those capabilities actually available in an implementation. - No file count, size distribution, memory limit, filesystem snapshot guarantee, or symbolic-link policy is supplied. ### Clarifying Questions to Ask - Can files change while they are being scanned, and can the helpers read a stable snapshot or version? - Are symbolic links followed, and are two hard-link paths to one physical file considered duplicates or one file? - Are empty files included, and how are unreadable files reported? - Does the result require exact equality, or may a cryptographic-digest match be accepted probabilistically? ### Part 1 — Reduce Candidate Comparisons Describe a staged approach that uses cheap information before reading complete file contents. Explain which filters can reject equality and which cannot prove it. #### What This Part Should Cover - Size grouping and optional partial-content filters. - Streaming full-content digests with bounded buffers. - Exact verification when the contract requires zero hash-collision risk. ### Part 2 — Make I/O and Results Reliable Explain bounded concurrency, file changes, read failures, grouping, and resource cleanup. #### What This Part Should Cover - No whole-file loading requirement for large files. - A consistency policy for files that mutate during scanning. - Error outcomes distinct from successfully inspected nonduplicates. ```hint Equal size is only a filter Different sizes prove that two byte sequences differ. Equal sizes, matching prefixes, or matching digests require different levels of evidence before declaring equality. ``` ### What a Strong Answer Covers - A complete content-equality grouping algorithm with controlled I/O. - Honest collision and concurrent-modification handling. - Complexity expressed in file count and bytes read, not only the number of filenames. ### Follow-up Questions - When could reading a small prefix reduce work, and when would it add an unnecessary extra read? - Why might more parallel readers hurt performance on one storage device? - What file-version evidence would make cached digests safe to reuse?

Overview: Find content-identical files with size filters, streaming digests, exact collision checks, bounded I/O, and explicit handling of changing or unreadable files.

Read the full Salient Ai Software Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/Salient Ai
Salient Ai logo
Salient Ai
Aug 3, 2026
hardSoftware EngineerOnsiteSoftware Engineering Fundamentals
0
0

Using supplied filesystem helpers, design an operation that finds files with identical contents. Explain how to reduce unnecessary file reads and memory use while preserving the distinction between exact content equality and a likely match based on metadata or hashing.

Constraints & Assumptions

  • The source names content-based duplicate-file detection and follow-ups on I/O and reading optimization. It does not supply the exact helper signatures.
  • Practice scope: report groups containing at least two paths whose complete byte contents are equal. Do not delete or overwrite files as part of duplicate detection.
  • Ask which helpers provide enumeration, size, bounded reads, and stable file identity or version. Use only those capabilities actually available in an implementation.
  • No file count, size distribution, memory limit, filesystem snapshot guarantee, or symbolic-link policy is supplied.

Clarifying Questions to Ask Guidance

  • Can files change while they are being scanned, and can the helpers read a stable snapshot or version?
  • Are symbolic links followed, and are two hard-link paths to one physical file considered duplicates or one file?
  • Are empty files included, and how are unreadable files reported?
  • Does the result require exact equality, or may a cryptographic-digest match be accepted probabilistically?

Part 1 — Reduce Candidate Comparisons

Describe a staged approach that uses cheap information before reading complete file contents. Explain which filters can reject equality and which cannot prove it.

What This Part Should Cover Guidance

  • Size grouping and optional partial-content filters.
  • Streaming full-content digests with bounded buffers.
  • Exact verification when the contract requires zero hash-collision risk.

Part 2 — Make I/O and Results Reliable

Explain bounded concurrency, file changes, read failures, grouping, and resource cleanup.

What This Part Should Cover Guidance

  • No whole-file loading requirement for large files.
  • A consistency policy for files that mutate during scanning.
  • Error outcomes distinct from successfully inspected nonduplicates.

What a Strong Answer Covers Guidance

  • A complete content-equality grouping algorithm with controlled I/O.
  • Honest collision and concurrent-modification handling.
  • Complexity expressed in file count and bytes read, not only the number of filenames.

Follow-up Questions Guidance

  • When could reading a small prefix reduce work, and when would it add an unnecessary extra read?
  • Why might more parallel readers hurt performance on one storage device?
  • What file-version evidence would make cached digests safe to reuse?
Loading comments...