Design a Photo Storage Service That Deduplicates Uploads by SHA-256 Content Hash
Company: OpenAI
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Technical Screen
Design a photo storage service in which users can upload photos, download them and delete them. Storage is deduplicated by content: if an uploaded photo's content has the same SHA-256 hash as a photo already stored, the system keeps only one copy of it.
```hint Two kinds of record
Separate what a user owns (their photo, its name, when they uploaded it) from the bytes that are stored once. Decide how the two are linked.
```
```hint Delete is the hard operation
A user deletes a photo whose bytes other photos also point to. Work out when the bytes may actually be removed, and what happens if another upload of the same content arrives at that moment.
```
```hint Who computes the hash
Consider which side computes the SHA-256, when the server can trust it, and whether the bytes need to be sent at all when the hash is already known.
```
### Constraints and Clarifications
- Deduplication is by exact content: two files are the same photo only if their SHA-256 hashes match. A resized or re-encoded copy is different content.
- Scale, photo sizes and latency targets are not given. State your assumptions.
### Clarifying Questions
- Is deduplication across all users, or only within one user's own library?
- If the same user uploads the same content twice, do they get two photos backed by one copy, or is the second upload rejected?
- When a user deletes a photo, must it disappear from their library immediately, and must the bytes be physically erased by some deadline once nobody references them?
- Are thumbnails or other derived versions in scope, and should they be deduplicated too?
- Can a user share a photo with others, or can only its uploader read it?
### What a Strong Answer Covers
- A data model that separates per-user photo records from content blobs keyed by SHA-256, with reference tracking
- An upload flow covering hashing, the duplicate check, and whether a known hash lets the client skip sending bytes
- A download flow that authorizes on the user's photo record rather than on the hash
- A delete flow and garbage collection of unreferenced blobs, including the race with a concurrent upload of the same content
- Metadata and object storage choices and how they scale, plus failure handling for partial uploads and crashed jobs
- The privacy and security consequences of deduplicating across users
### Follow-up Questions
- An attacker knows the SHA-256 of a photo they suspect another user has stored. What can they learn or obtain through your upload API, and how do you prevent it?
- How would you detect and repair reference counts that drift from the real number of photos pointing at a blob?
- A user uploads a large photo over an unreliable mobile connection. How does the upload survive interruptions without storing a partial file as a blob?
- How would you extend deduplication to thumbnails, and is it worth it?
Overview: A system design question about a photo service that supports upload, download and delete while keeping only one stored copy of any content with the same SHA-256 hash. It tests content-addressed storage, reference tracking, safe deletion under concurrent uploads, and the security risks of deduplicating across users.
Read the full OpenAI Software Engineer interview experience this question came from