Docker optimization, container fundamentals, and CI/CD for an ML model service

Read the full interview experience this question came from →

Quick Overview

Practice explaining how Docker images, layers, build caching and container isolation work, then optimizing the image and runtime of a Python ML inference service. The question also covers continuous integration, delivery and deployment for shipping a containerized model safely.

Docker optimization, container fundamentals, and CI/CD for an ML model service

Company: eBay

Role: Applied Scientist

Category: Software Engineering Fundamentals

Difficulty: easy

Interview Round: Technical Screen

You maintain a Python machine learning inference service that is packaged as a Docker image and shipped to production through an automated pipeline. The discussion covers three connected topics: how Docker works underneath, how you would optimize Docker for this service, and the concepts behind CI/CD for delivering it. ### Constraints and Clarifications - Assume the service loads a trained model artifact and serves predictions over HTTP, with typical ML dependencies (a deep learning framework, tokenization or feature-processing libraries, native extensions). - No size, build-time, startup, or latency targets were given. Treat the target as something to clarify, and justify each optimization by the bottleneck it removes. ### Clarifying Questions - Which cost hurts most today: image size and pull time, CI build time, container cold start, or steady-state inference throughput? - Does the service run on CPU hosts or GPU hosts? - Are the model weights baked into the image, or fetched at startup from a model registry or object store? - What runs the containers (single hosts or an orchestrator such as Kubernetes), and how often are images rebuilt and redeployed? ### Part 1 — How Docker works Explain the difference between an image and a container, how image layers and the build cache work, and what isolates a container from the host. Contrast a container with a virtual machine. ```hint Follow one changed line Trace what happens to every later instruction in a Dockerfile when one earlier instruction changes. ``` #### What This Part Should Cover - Image versus container: read-only layers plus a writable container layer - Build cache behavior and why instruction order matters - Kernel-level isolation and resource limits versus hardware virtualization, with the practical consequences ### Part 2 — Optimize the service's image and container Start from this representative Dockerfile and explain how you would make the image smaller, faster to build, faster to start, and efficient at runtime. ```dockerfile FROM python:3.11 WORKDIR /app COPY . . RUN apt-get update && apt-get install -y build-essential git RUN pip install -r requirements.txt CMD ["python", "serve.py"] ``` ```hint Build versus run List what the build needs and what the running container needs; they are not the same set. ``` ```hint Change frequency Order the steps by how often their inputs change. ``` #### What This Part Should Cover - Image size: base image choice, multi-stage builds, removing build tools and caches, a `.dockerignore` - Build speed: layer ordering, dependency caching, cache reuse in CI - Startup and runtime: where the model comes from, worker and thread settings versus container CPU and memory limits, health checks - Security and reproducibility: pinned versions and a non-root user ### Part 3 — CI/CD for the containerized service Explain continuous integration, continuous delivery, and continuous deployment, then describe a pipeline that takes a change (code or model) from commit to a running production container for this service. ```hint Gates Decide what must block a merge and what must block a production rollout. ``` #### What This Part Should Cover - The three terms and the difference between delivery and deployment - Pipeline stages from tests to an immutable, versioned image in a registry - ML-specific checks and how the model version is tied to what is deployed - Safe rollout and rollback ### What a Strong Answer Covers - Ties every optimization to a measured bottleneck instead of reciting a checklist - Reasons from Docker internals (layers, cache, isolation) to concrete Dockerfile and runtime changes - Treats the image as an immutable, versioned artifact from CI through production - Handles ML-specific concerns: large weights, native dependencies, CPU versus GPU base images, and model versioning ### Follow-up Questions - Your container is killed for exceeding its memory limit under load even though the host has free memory. How do you debug it? - How do you decide between baking model weights into the image and downloading them at startup, and how does that choice affect rollback? - Two builds from the same commit behave differently in production. How do you make builds reproducible? - A new image passes every CI check but degrades prediction quality after rollout. What should the pipeline have caught, and how do you roll back?

Overview: Practice explaining how Docker images, layers, build caching and container isolation work, then optimizing the image and runtime of a Python ML inference service. The question also covers continuous integration, delivery and deployment for shipping a containerized model safely.

Read the full eBay Applied Scientist interview experience this question came from

|Home/Software Engineering Fundamentals/eBay
eBay logo
eBay
Sep 24, 2026
easyApplied ScientistTechnical ScreenSoftware Engineering Fundamentals
0
0

You maintain a Python machine learning inference service that is packaged as a Docker image and shipped to production through an automated pipeline. The discussion covers three connected topics: how Docker works underneath, how you would optimize Docker for this service, and the concepts behind CI/CD for delivering it.

Constraints and Clarifications

  • Assume the service loads a trained model artifact and serves predictions over HTTP, with typical ML dependencies (a deep learning framework, tokenization or feature-processing libraries, native extensions).
  • No size, build-time, startup, or latency targets were given. Treat the target as something to clarify, and justify each optimization by the bottleneck it removes.

Clarifying Questions Guidance

  • Which cost hurts most today: image size and pull time, CI build time, container cold start, or steady-state inference throughput?
  • Does the service run on CPU hosts or GPU hosts?
  • Are the model weights baked into the image, or fetched at startup from a model registry or object store?
  • What runs the containers (single hosts or an orchestrator such as Kubernetes), and how often are images rebuilt and redeployed?

Part 1 — How Docker works

Explain the difference between an image and a container, how image layers and the build cache work, and what isolates a container from the host. Contrast a container with a virtual machine.

What This Part Should Cover Guidance

  • Image versus container: read-only layers plus a writable container layer
  • Build cache behavior and why instruction order matters
  • Kernel-level isolation and resource limits versus hardware virtualization, with the practical consequences

Part 2 — Optimize the service's image and container

Start from this representative Dockerfile and explain how you would make the image smaller, faster to build, faster to start, and efficient at runtime.

FROM python:3.11
WORKDIR /app
COPY . .
RUN apt-get update && apt-get install -y build-essential git
RUN pip install -r requirements.txt
CMD ["python", "serve.py"]

What This Part Should Cover Guidance

  • Image size: base image choice, multi-stage builds, removing build tools and caches, a .dockerignore
  • Build speed: layer ordering, dependency caching, cache reuse in CI
  • Startup and runtime: where the model comes from, worker and thread settings versus container CPU and memory limits, health checks
  • Security and reproducibility: pinned versions and a non-root user

Part 3 — CI/CD for the containerized service

Explain continuous integration, continuous delivery, and continuous deployment, then describe a pipeline that takes a change (code or model) from commit to a running production container for this service.

What This Part Should Cover Guidance

  • The three terms and the difference between delivery and deployment
  • Pipeline stages from tests to an immutable, versioned image in a registry
  • ML-specific checks and how the model version is tied to what is deployed
  • Safe rollout and rollback

What a Strong Answer Covers Guidance

  • Ties every optimization to a measured bottleneck instead of reciting a checklist
  • Reasons from Docker internals (layers, cache, isolation) to concrete Dockerfile and runtime changes
  • Treats the image as an immutable, versioned artifact from CI through production
  • Handles ML-specific concerns: large weights, native dependencies, CPU versus GPU base images, and model versioning

Follow-up Questions Guidance

  • Your container is killed for exceeding its memory limit under load even though the host has free memory. How do you debug it?
  • How do you decide between baking model weights into the image and downloading them at startup, and how does that choice affect rollback?
  • Two builds from the same commit behave differently in production. How do you make builds reproducible?
  • A new image passes every CI check but degrades prediction quality after rollout. What should the pipeline have caught, and how do you roll back?
Loading comments...