Design a Parallel and Resumable Model Downloader

Read the full interview experience this question came from →

Quick Overview

Design a parallel model downloader with version-bound byte ranges, resumable verified pieces, bounded resource use, and atomic publication of complete artifacts.

Design a Parallel and Resumable Model Downloader

Company: Anthropic

Role: Software Engineer

Category: System Design

Difficulty: medium

Interview Round: Onsite

Design a model downloader that downloads a model artifact in parallel pieces, resumes interrupted work, and makes a complete model available to its consumer. ### Constraints & Assumptions - The source is a compilation of third-party interview reports and names sharded model downloading as a system-design task. It supplies no proprietary architecture, measured performance, or detailed original interface. - **Practice scope:** a published model version has an immutable manifest listing its files, expected lengths, and integrity digests. Downloads may use file-level parallelism and byte ranges when the server supports them. The consumer must not mistake partial files for a complete model. - Network requests and the downloader can fail. A retry should reuse verified work when its artifact binding still matches. - No artifact size, concurrency limit, bandwidth target, or distribution topology is supplied; explain how these inputs determine the design. ### Clarifying Questions to Ask - Is the source content immutable or versioned, and does it support byte-range requests with version validation? - Are integrity digests available for complete files, individual chunks, or both? - Must several processes or machines share a local cache, and who owns an in-progress download? - Does the consumer require atomic publication of the complete model version or only independent files? ### Part 1 — Plan and Download Pieces Describe the manifest, piece plan, concurrency control, and transfer validation. Explain how to avoid combining ranges from different remote object versions. #### What This Part Should Cover - Stable artifact identity and explicit byte ranges. - Bounded network, disk, and memory use. - Verification of status, returned ranges, content length, and object version. ### Part 2 — Resume and Publish Safely Explain persistent progress, restart behavior, integrity checks, and the final transition from partial download to usable model. #### What This Part Should Cover - Durable verified-piece records bound to the correct manifest and object version. - Safe handling of incomplete writes or stale progress metadata. - Final validation and atomic visibility at the granularity the consumer requires. ```hint Bind progress to content identity The fact that bytes 0 through 999 were downloaded yesterday is useful only if they still belong to the exact artifact version being assembled today. ``` ### What a Strong Answer Covers - Parallel downloading that respects the source server's actual capabilities. - Resume correctness, content integrity, and controlled resource use. - A publication boundary that prevents consumers from loading partial or mixed-version model files. ### Follow-up Questions - What happens if the server ignores a range request and returns the entire file? - How would two local processes avoid independently writing the same partial model? - When can a completed cache entry be evicted without disrupting an active reader?

Overview: Design a parallel model downloader with version-bound byte ranges, resumable verified pieces, bounded resource use, and atomic publication of complete artifacts.

Read the full Anthropic Software Engineer interview experience this question came from

|Home/System Design/Anthropic
Anthropic logo
Anthropic
Oct 5, 2026
mediumSoftware EngineerOnsiteSystem Design
0
0

Design a model downloader that downloads a model artifact in parallel pieces, resumes interrupted work, and makes a complete model available to its consumer.

Constraints & Assumptions

  • The source is a compilation of third-party interview reports and names sharded model downloading as a system-design task. It supplies no proprietary architecture, measured performance, or detailed original interface.
  • Practice scope: a published model version has an immutable manifest listing its files, expected lengths, and integrity digests. Downloads may use file-level parallelism and byte ranges when the server supports them. The consumer must not mistake partial files for a complete model.
  • Network requests and the downloader can fail. A retry should reuse verified work when its artifact binding still matches.
  • No artifact size, concurrency limit, bandwidth target, or distribution topology is supplied; explain how these inputs determine the design.

Clarifying Questions to Ask Guidance

  • Is the source content immutable or versioned, and does it support byte-range requests with version validation?
  • Are integrity digests available for complete files, individual chunks, or both?
  • Must several processes or machines share a local cache, and who owns an in-progress download?
  • Does the consumer require atomic publication of the complete model version or only independent files?

Part 1 — Plan and Download Pieces

Describe the manifest, piece plan, concurrency control, and transfer validation. Explain how to avoid combining ranges from different remote object versions.

What This Part Should Cover Guidance

  • Stable artifact identity and explicit byte ranges.
  • Bounded network, disk, and memory use.
  • Verification of status, returned ranges, content length, and object version.

Part 2 — Resume and Publish Safely

Explain persistent progress, restart behavior, integrity checks, and the final transition from partial download to usable model.

What This Part Should Cover Guidance

  • Durable verified-piece records bound to the correct manifest and object version.
  • Safe handling of incomplete writes or stale progress metadata.
  • Final validation and atomic visibility at the granularity the consumer requires.

What a Strong Answer Covers Guidance

  • Parallel downloading that respects the source server's actual capabilities.
  • Resume correctness, content integrity, and controlled resource use.
  • A publication boundary that prevents consumers from loading partial or mixed-version model files.

Follow-up Questions Guidance

  • What happens if the server ignores a range request and returns the entire file?
  • How would two local processes avoid independently writing the same partial model?
  • When can a completed cache entry be evicted without disrupting an active reader?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...