Find Duplicate Files by Content with Efficient I/O
Company: Salient Ai
Role: Software Engineer
Category: Software Engineering Fundamentals
Difficulty: hard
Interview Round: Onsite
Using supplied filesystem helpers, design an operation that finds files with identical contents. Explain how to reduce unnecessary file reads and memory use while preserving the distinction between exact content equality and a likely match based on metadata or hashing.
### Constraints & Assumptions
- The source names content-based duplicate-file detection and follow-ups on I/O and reading optimization. It does not supply the exact helper signatures.
- **Practice scope:** report groups containing at least two paths whose complete byte contents are equal. Do not delete or overwrite files as part of duplicate detection.
- Ask which helpers provide enumeration, size, bounded reads, and stable file identity or version. Use only those capabilities actually available in an implementation.
- No file count, size distribution, memory limit, filesystem snapshot guarantee, or symbolic-link policy is supplied.
### Clarifying Questions to Ask
- Can files change while they are being scanned, and can the helpers read a stable snapshot or version?
- Are symbolic links followed, and are two hard-link paths to one physical file considered duplicates or one file?
- Are empty files included, and how are unreadable files reported?
- Does the result require exact equality, or may a cryptographic-digest match be accepted probabilistically?
### Part 1 — Reduce Candidate Comparisons
Describe a staged approach that uses cheap information before reading complete file contents. Explain which filters can reject equality and which cannot prove it.
#### What This Part Should Cover
- Size grouping and optional partial-content filters.
- Streaming full-content digests with bounded buffers.
- Exact verification when the contract requires zero hash-collision risk.
### Part 2 — Make I/O and Results Reliable
Explain bounded concurrency, file changes, read failures, grouping, and resource cleanup.
#### What This Part Should Cover
- No whole-file loading requirement for large files.
- A consistency policy for files that mutate during scanning.
- Error outcomes distinct from successfully inspected nonduplicates.
```hint Equal size is only a filter
Different sizes prove that two byte sequences differ. Equal sizes, matching prefixes, or matching digests require different levels of evidence before declaring equality.
```
### What a Strong Answer Covers
- A complete content-equality grouping algorithm with controlled I/O.
- Honest collision and concurrent-modification handling.
- Complexity expressed in file count and bytes read, not only the number of filenames.
### Follow-up Questions
- When could reading a small prefix reduce work, and when would it add an unnecessary extra read?
- Why might more parallel readers hurt performance on one storage device?
- What file-version evidence would make cached digests safe to reuse?
Overview: Find content-identical files with size filters, streaming digests, exact collision checks, bounded I/O, and explicit handling of changing or unreadable files.
Read the full Salient Ai Software Engineer interview experience this question came from