Design a Video Upload Service for 100 GB Files with Chunked, Resumable Uploads

Quick Overview

A system design question about building a service that accepts video uploads of up to 100 GB. It covers moving very large files efficiently, splitting them into chunks, resuming interrupted uploads, choosing storage for bytes and metadata, and scaling to many concurrent uploads.

Design a Video Upload Service for 100 GB Files with Chunked, Resumable Uploads

Company: Oracle

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Technical Screen

Design a system that lets users upload video files as large as 100 GB. The interviewer steered the discussion through five areas, in this order: handling very large uploads efficiently, chunking, resumable uploads, storage choices, and scalability. Work through each one. ### Clarifying Questions - Which clients upload (browsers, mobile apps or desktop tools), and over what kinds of network? This affects chunk size and how often an upload is interrupted. - Does the scope end once the original file is stored durably, or does it include transcoding and playback? - How many uploads run at the same time, and how much data arrives per day? - Must the system verify the integrity of the whole file end to end, and should a file that was already uploaded be detected instead of sent again? - How long may an unfinished upload wait before its partial data is deleted? ### Part 1 — Handling a very large upload efficiently Explain why a 100 GB file should not be sent as one HTTP request through the application servers, and describe the path the bytes should take instead. ```hint Follow the bytes Trace where each byte goes in one long request through your API servers, and ask which components actually need to see the video data and which only need to know about it. ``` #### What This Part Should Cover - The concrete failure modes of one long request at this size: timeouts, server memory and disk, and retries from zero - Separating the control path (authorization, metadata) from the data path (the bytes) - Transfer time for 100 GB at realistic bandwidths, and the client-side levers for throughput ### Part 2 — Chunking Describe how the file is split into chunks, how the chunk size is chosen, how the chunks are uploaded, and how they become one video at the end. ```hint Size the parts Relate the chunk size to the number of chunks for a 100 GB file, to the data re-sent when one chunk fails, and to any per-upload limit on the number of parts that your storage imposes. ``` #### What This Part Should Cover - Chunk-size arithmetic for 100 GB and the trade-off it balances - Parallel chunk uploads, per-chunk integrity checks and retries - The commit step that assembles the chunks into one stored object ### Part 3 — Resumable uploads The network drops partway through, or the client restarts. How does the upload continue from where it stopped instead of starting over? ```hint Who remembers progress Decide where the authoritative record of which chunks have arrived lives, and how a client that lost its local state finds its unfinished upload again. ``` #### What This Part Should Cover - Upload-session state: what it holds, where it lives and when it expires - The resume handshake that tells the client which chunks are missing - Idempotent retries, including a chunk that arrived but whose acknowledgment was lost ### Part 4 — Storage Choose where the video bytes, the in-progress chunks and the upload metadata live. ```hint Bytes versus records Separate the large, write-once bytes from the small, frequently updated records, and pick a store for each by its access pattern. ``` #### What This Part Should Cover - Object storage for the video data versus a database for metadata and session state - Durability, cost tiers and lifecycle rules for the original files - Reclaiming the space held by abandoned partial uploads ### Part 5 — Scalability The service grows to many concurrent uploads from users around the world. What scales, and how? ```hint Find what grows Identify which components see load that grows with the number of bytes uploaded and which see load that grows only with the number of requests, then scale each kind separately. ``` #### What This Part Should Cover - A stateless API tier that never handles video bytes - Ingest close to users, plus per-user limits and quotas - Asynchronous work after upload with backpressure, and avoiding hot spots in the metadata store ### What a Strong Answer Covers - A clear split between the control plane and the data plane, carried consistently through every part - Numbers worked out for a 100 GB file rather than hand-waving - Failure handling end to end: retries, idempotency, integrity checks, expiry and cleanup - Security of the upload path: short-lived, narrowly scoped upload grants and size limits - Observability: the metrics and alerts that show whether uploads are succeeding ### Follow-up Questions - How would you avoid re-sending data when a user uploads a file the system already has? - How do you stop a client from filling storage with uploads it never completes? - How could processing such as transcoding start before the whole file has arrived? - How would you show the user accurate upload progress, including right after a resume?

Overview: A system design question about building a service that accepts video uploads of up to 100 GB. It covers moving very large files efficiently, splitting them into chunks, resuming interrupted uploads, choosing storage for bytes and metadata, and scaling to many concurrent uploads.

|Home/System Design/Oracle
Oracle logo
Oracle
Sep 14, 2026
mediumSoftware EngineerTechnical ScreenSystem Design
0
0

Design a system that lets users upload video files as large as 100 GB. The interviewer steered the discussion through five areas, in this order: handling very large uploads efficiently, chunking, resumable uploads, storage choices, and scalability. Work through each one.

Clarifying Questions Guidance

  • Which clients upload (browsers, mobile apps or desktop tools), and over what kinds of network? This affects chunk size and how often an upload is interrupted.
  • Does the scope end once the original file is stored durably, or does it include transcoding and playback?
  • How many uploads run at the same time, and how much data arrives per day?
  • Must the system verify the integrity of the whole file end to end, and should a file that was already uploaded be detected instead of sent again?
  • How long may an unfinished upload wait before its partial data is deleted?

Part 1 — Handling a very large upload efficiently

Explain why a 100 GB file should not be sent as one HTTP request through the application servers, and describe the path the bytes should take instead.

What This Part Should Cover Guidance

  • The concrete failure modes of one long request at this size: timeouts, server memory and disk, and retries from zero
  • Separating the control path (authorization, metadata) from the data path (the bytes)
  • Transfer time for 100 GB at realistic bandwidths, and the client-side levers for throughput

Part 2 — Chunking

Describe how the file is split into chunks, how the chunk size is chosen, how the chunks are uploaded, and how they become one video at the end.

What This Part Should Cover Guidance

  • Chunk-size arithmetic for 100 GB and the trade-off it balances
  • Parallel chunk uploads, per-chunk integrity checks and retries
  • The commit step that assembles the chunks into one stored object

Part 3 — Resumable uploads

The network drops partway through, or the client restarts. How does the upload continue from where it stopped instead of starting over?

What This Part Should Cover Guidance

  • Upload-session state: what it holds, where it lives and when it expires
  • The resume handshake that tells the client which chunks are missing
  • Idempotent retries, including a chunk that arrived but whose acknowledgment was lost

Part 4 — Storage

Choose where the video bytes, the in-progress chunks and the upload metadata live.

What This Part Should Cover Guidance

  • Object storage for the video data versus a database for metadata and session state
  • Durability, cost tiers and lifecycle rules for the original files
  • Reclaiming the space held by abandoned partial uploads

Part 5 — Scalability

The service grows to many concurrent uploads from users around the world. What scales, and how?

What This Part Should Cover Guidance

  • A stateless API tier that never handles video bytes
  • Ingest close to users, plus per-user limits and quotas
  • Asynchronous work after upload with backpressure, and avoiding hot spots in the metadata store

What a Strong Answer Covers Guidance

  • A clear split between the control plane and the data plane, carried consistently through every part
  • Numbers worked out for a 100 GB file rather than hand-waving
  • Failure handling end to end: retries, idempotency, integrity checks, expiry and cleanup
  • Security of the upload path: short-lived, narrowly scoped upload grants and size limits
  • Observability: the metrics and alerts that show whether uploads are succeeding

Follow-up Questions Guidance

  • How would you avoid re-sending data when a user uploads a file the system already has?
  • How do you stop a client from filling storage with uploads it never completes?
  • How could processing such as transcoding start before the whole file has arrived?
  • How would you show the user accurate upload progress, including right after a resume?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...