Docker optimization, container fundamentals, and CI/CD for an ML model service
Company: eBay
Role: Applied Scientist
Category: Software Engineering Fundamentals
Difficulty: easy
Interview Round: Technical Screen
You maintain a Python machine learning inference service that is packaged as a Docker image and shipped to production through an automated pipeline. The discussion covers three connected topics: how Docker works underneath, how you would optimize Docker for this service, and the concepts behind CI/CD for delivering it.
### Constraints and Clarifications
- Assume the service loads a trained model artifact and serves predictions over HTTP, with typical ML dependencies (a deep learning framework, tokenization or feature-processing libraries, native extensions).
- No size, build-time, startup, or latency targets were given. Treat the target as something to clarify, and justify each optimization by the bottleneck it removes.
### Clarifying Questions
- Which cost hurts most today: image size and pull time, CI build time, container cold start, or steady-state inference throughput?
- Does the service run on CPU hosts or GPU hosts?
- Are the model weights baked into the image, or fetched at startup from a model registry or object store?
- What runs the containers (single hosts or an orchestrator such as Kubernetes), and how often are images rebuilt and redeployed?
### Part 1 — How Docker works
Explain the difference between an image and a container, how image layers and the build cache work, and what isolates a container from the host. Contrast a container with a virtual machine.
```hint Follow one changed line
Trace what happens to every later instruction in a Dockerfile when one earlier instruction changes.
```
#### What This Part Should Cover
- Image versus container: read-only layers plus a writable container layer
- Build cache behavior and why instruction order matters
- Kernel-level isolation and resource limits versus hardware virtualization, with the practical consequences
### Part 2 — Optimize the service's image and container
Start from this representative Dockerfile and explain how you would make the image smaller, faster to build, faster to start, and efficient at runtime.
```dockerfile
FROM python:3.11
WORKDIR /app
COPY . .
RUN apt-get update && apt-get install -y build-essential git
RUN pip install -r requirements.txt
CMD ["python", "serve.py"]
```
```hint Build versus run
List what the build needs and what the running container needs; they are not the same set.
```
```hint Change frequency
Order the steps by how often their inputs change.
```
#### What This Part Should Cover
- Image size: base image choice, multi-stage builds, removing build tools and caches, a `.dockerignore`
- Build speed: layer ordering, dependency caching, cache reuse in CI
- Startup and runtime: where the model comes from, worker and thread settings versus container CPU and memory limits, health checks
- Security and reproducibility: pinned versions and a non-root user
### Part 3 — CI/CD for the containerized service
Explain continuous integration, continuous delivery, and continuous deployment, then describe a pipeline that takes a change (code or model) from commit to a running production container for this service.
```hint Gates
Decide what must block a merge and what must block a production rollout.
```
#### What This Part Should Cover
- The three terms and the difference between delivery and deployment
- Pipeline stages from tests to an immutable, versioned image in a registry
- ML-specific checks and how the model version is tied to what is deployed
- Safe rollout and rollback
### What a Strong Answer Covers
- Ties every optimization to a measured bottleneck instead of reciting a checklist
- Reasons from Docker internals (layers, cache, isolation) to concrete Dockerfile and runtime changes
- Treats the image as an immutable, versioned artifact from CI through production
- Handles ML-specific concerns: large weights, native dependencies, CPU versus GPU base images, and model versioning
### Follow-up Questions
- Your container is killed for exceeding its memory limit under load even though the host has free memory. How do you debug it?
- How do you decide between baking model weights into the image and downloading them at startup, and how does that choice affect rollback?
- Two builds from the same commit behave differently in production. How do you make builds reproducible?
- A new image passes every CI check but degrades prediction quality after rollout. What should the pipeline have caught, and how do you roll back?
Overview: Practice explaining how Docker images, layers, build caching and container isolation work, then optimizing the image and runtime of a Python ML inference service. The question also covers continuous integration, delivery and deployment for shipping a containerized model safely.
Read the full eBay Applied Scientist interview experience this question came from