Build a Resize-and-Rotate Image Pipeline with Multiprocessing

Read the full interview experience this question came from →

Quick Overview

Design a Pillow resize-and-rotate pipeline with explicit transform semantics, bounded multiprocessing, per-image failures, and safe output publication.

Build a Resize-and-Rotate Image Pipeline with Multiprocessing

Company: Anthropic

Role: Software Engineer

Category: Software Engineering Fundamentals

Difficulty: medium

Interview Round: Onsite

Design an image-processing pipeline that uses Pillow to resize and rotate images, then extend it to process independent images using multiple processes. ### Constraints & Assumptions - The source specifically names resize/rotate operations with the image library and a multiprocessing extension; it does not supply a complete original API or transform specification. - **Practice scope:** each input image has an explicit transform specification describing target size, rotation angle, operation order, and output destination. Produce one result or error record per input. Do not invent a hidden fixed order or angle. - Input images are independent, while disk bandwidth and memory are shared resources. - Define whether rotation expands the output canvas, how resize handles aspect ratio, and which output format is required before evaluating image dimensions or appearance. - No exact process count, file size, throughput target, or acceptable lossy encoding is supplied. ### Clarifying Questions to Ask - Is resize a stretch, aspect-ratio-preserving fit, or crop, and which resampling mode is required? - Does rotation preserve the original canvas or expand it, and does it happen before or after resize? - Can two inputs target the same output path, and may an existing output be replaced? - Should one invalid image abort the batch or be reported while other images continue? ### Part 1 — Build a Correct Single-Image Operation Describe opening and decoding the image, applying the specified transformations, writing the result, and recording failure. Identify where format, orientation, and resource lifecycle matter. #### What This Part Should Cover - Explicit transform order and dimensions without guessing a missing source rule. - Proper ownership and closure of image/file resources. - Output publication that does not confuse a truncated file with a completed result. ### Part 2 — Add Multiprocessing Explain how to distribute independent images across a bounded process set. Describe what is passed to a worker, how results are collected, and how failures and resource pressure are handled. #### What This Part Should Cover - Small serializable work descriptions rather than unnecessary transfer of decoded pixel buffers. - Bounded in-flight work and memory-aware concurrency. - Per-input result identity, worker failure handling, and cleanup. ```hint Account for decoded size A compressed file's size is not the memory occupied by its decoded pixels or the temporary images created during transformation. ``` ### What a Strong Answer Covers - Correct library-based transformation under an explicit image contract. - Process isolation and bounded parallel execution without shared-output races. - Per-image failure reporting, resource cleanup, and evidence-based throughput reasoning. ### Follow-up Questions - Why can increasing process count make a disk-bound batch slower? - What changes when resize and rotation are applied in the opposite order? - How would a retry avoid publishing two conflicting results for the same destination?

Overview: Design a Pillow resize-and-rotate pipeline with explicit transform semantics, bounded multiprocessing, per-image failures, and safe output publication.

Read the full Anthropic Software Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/Anthropic
Anthropic logo
Anthropic
Oct 5, 2026
mediumSoftware EngineerOnsiteSoftware Engineering Fundamentals
0
0

Design an image-processing pipeline that uses Pillow to resize and rotate images, then extend it to process independent images using multiple processes.

Constraints & Assumptions

  • The source specifically names resize/rotate operations with the image library and a multiprocessing extension; it does not supply a complete original API or transform specification.
  • Practice scope: each input image has an explicit transform specification describing target size, rotation angle, operation order, and output destination. Produce one result or error record per input. Do not invent a hidden fixed order or angle.
  • Input images are independent, while disk bandwidth and memory are shared resources.
  • Define whether rotation expands the output canvas, how resize handles aspect ratio, and which output format is required before evaluating image dimensions or appearance.
  • No exact process count, file size, throughput target, or acceptable lossy encoding is supplied.

Clarifying Questions to Ask Guidance

  • Is resize a stretch, aspect-ratio-preserving fit, or crop, and which resampling mode is required?
  • Does rotation preserve the original canvas or expand it, and does it happen before or after resize?
  • Can two inputs target the same output path, and may an existing output be replaced?
  • Should one invalid image abort the batch or be reported while other images continue?

Part 1 — Build a Correct Single-Image Operation

Describe opening and decoding the image, applying the specified transformations, writing the result, and recording failure. Identify where format, orientation, and resource lifecycle matter.

What This Part Should Cover Guidance

  • Explicit transform order and dimensions without guessing a missing source rule.
  • Proper ownership and closure of image/file resources.
  • Output publication that does not confuse a truncated file with a completed result.

Part 2 — Add Multiprocessing

Explain how to distribute independent images across a bounded process set. Describe what is passed to a worker, how results are collected, and how failures and resource pressure are handled.

What This Part Should Cover Guidance

  • Small serializable work descriptions rather than unnecessary transfer of decoded pixel buffers.
  • Bounded in-flight work and memory-aware concurrency.
  • Per-input result identity, worker failure handling, and cleanup.

What a Strong Answer Covers Guidance

  • Correct library-based transformation under an explicit image contract.
  • Process isolation and bounded parallel execution without shared-output races.
  • Per-image failure reporting, resource cleanup, and evidence-based throughput reasoning.

Follow-up Questions Guidance

  • Why can increasing process count make a disk-bound batch slower?
  • What changes when resize and rotation are applied in the opposite order?
  • How would a retry avoid publishing two conflicting results for the same destination?
Loading comments...