DevOps Engineer Interview Questions 2026: CI/CD, Docker, Kubernetes, Terraform, and Incidents

Prepare for DevOps engineer interviews with CI/CD, Docker, Kubernetes, Terraform, incident response, debugging, and production trade-offs.

Author: PracHub

Published: 9/2/2026

DevOps Engineer Interview Questions 2026: CI/CD, Docker, Kubernetes, Terraform, and Incidents

September 2, 2026

Quick Overview

Prepare for DevOps engineer interviews with practical coverage of CI/CD, Docker, Kubernetes, Terraform, observability, incident response, and trade-off reasoning.

DevOps EngineerFree

A strong DevOps engineer interview answer connects code change, infrastructure state, service health, and recovery. Expect questions about CI/CD, Docker, Kubernetes, Terraform, observability, and incidents—but the real test is whether you can make a risky change traceable, limit its blast radius, diagnose failure from evidence, and restore service safely.

DevOps engineer interview: an evaluation of whether a candidate can automate reliable software delivery, operate infrastructure, troubleshoot across system boundaries, and communicate production trade-offs. The exact rounds and tooling vary by company, team, and level.

This guide is for engineers preparing for delivery, infrastructure, troubleshooting, design, and behavioral rounds. Start with PracHub DevOps Engineer interview questions, then use the deeper tool-specific guides below for targeted practice.

DevOps engineer interview preparation across CI CD containers Kubernetes Terraform observability and incidents

The strongest preparation follows one change from source and artifact through infrastructure, runtime health, telemetry, and recovery.

DevOps engineer interview questions: the quick answer

Prepare six connected areas: CI/CD, containers, Kubernetes, infrastructure as code, observability, and incident response. In every answer, make four things explicit: the desired outcome, the state that must remain correct, the evidence you would inspect, and the safest recovery path.

Use the CHANGE model for delivery and infrastructure questions:

  1. Clarify the outcome, constraints, and acceptable risk.
  2. Hold one immutable, traceable artifact.
  3. Automate tests, policy, security, and approval evidence.
  4. Narrow exposure with staged rollout and blast-radius controls.
  5. Guard service health with user-facing operational signals.
  6. Exit safely through rollback, roll-forward, or containment.

CHANGE is a PracHub preparation model, not a universal employer rubric.

What interviewers evaluate beyond tool knowledge

DevOps interviews often begin with a tool, then move quickly into ownership and failure. “I would use Kubernetes” is incomplete until you explain which Kubernetes behavior solves the requirement, how it fails, and what proves recovery.

SignalStrong answerWeak answer
Delivery safetyPromotes a traceable artifact through explicit gatesRebuilds separately in every environment
Systems reasoningFollows state and requests across boundariesNames components without contracts
Automation judgmentAutomates repeatable policy and keeps exceptions visibleAutomates a fragile process without safeguards
TroubleshootingForms ranked hypotheses from customer impact and evidenceRestarts components at random
SecurityUses least privilege, short-lived identity, provenance, and auditabilityTreats secrets scanning as the whole security model
CommunicationStates assumptions, owners, mitigation, and verificationGives commands without explaining the decision

The expected radius grows with seniority. Entry-level candidates may diagnose one failing container; senior candidates should connect organization-wide delivery controls, multi-service failure, migration risk, and incident coordination.

CI/CD interview questions

How would you design a pipeline from commit to production?

Begin with the artifact, not a list of stages. Build once in an isolated environment, attach the source revision and dependency information, run fast feedback first, and promote the same immutable artifact through later environments. Define what each gate proves and who can approve an exception.

GitHub deployment environments illustrate useful controls: approval rules, branch restrictions, and environment-scoped secrets. Artifact attestations add provenance. In an interview, connect those mechanisms to the threat or failure they contain.

When should a deployment roll back automatically?

Automatic rollback is appropriate when the signal is trustworthy, the observation window is meaningful, and the previous version remains compatible with current data and dependencies. It is dangerous after an irreversible migration or when noisy metrics oscillate between versions.

State the decision rule: stop exposure on a clear service-level regression; roll back only when compatibility is known; otherwise contain traffic or roll forward. A deployment is not healthy merely because the workflow completed or containers became ready.

How do you handle flaky tests without normalizing failure?

Preserve the first failure, use bounded reruns to collect evidence, track flake rate and ownership, and quarantine only with replacement coverage and a repair deadline. Hidden automatic retries turn a red signal green without making the system more reliable.

For deeper scenarios on artifacts, migrations, secrets, and rollback, use CI/CD Interview Questions for Senior Engineers.

Docker interview questions

How do you make a Docker image smaller and safer?

Use a small trusted base, multi-stage builds, a narrow build context, deliberate layer order, and a non-root runtime user where the application permits it. Remove compilers and package caches from the runtime image. Pinning a base by digest improves reproducibility, but it also creates an explicit update responsibility.

Docker’s build guidance recommends multi-stage builds, .dockerignore, cache-aware layering, trusted minimal bases, and regular rebuilds. A strong answer weighs size, build speed, debuggability, provenance, and patching instead of optimizing one number.

A container exits immediately. What do you inspect?

Start with the exit code, configured entrypoint and command, current and previous logs, environment, mounts, working directory, user, and resource limits. Confirm whether the application finished normally, failed startup, received a signal, or was killed by the platform.

Do not repeatedly restart before preserving evidence. Also separate image problems from runtime configuration, host pressure, and orchestrator behavior. The Docker interview guide covers image, networking, storage, and debugging scenarios in more depth.

Kubernetes interview questions

A Pod is Running, but users receive 503s. Where do you start?

Running describes a Pod phase, not successful requests. Check readiness, EndpointSlices, the Service selector and target port, ingress or gateway behavior, network policy, and application logs. Compare the failing user path with direct Pod and Service access to find the first broken contract.

Kubernetes probe documentation distinguishes startup, readiness, and liveness. A failed readiness probe removes a Pod from eligible Service endpoints; an incorrect liveness probe can create restart cascades.

How do you diagnose CrashLoopBackOff?

Inspect container state, reason, restart count, events, current logs, and --previous logs. Then test configuration, dependencies, permissions, filesystem assumptions, probes, and memory limits. CrashLoopBackOff is a symptom of repeated failure with backoff, not a root cause.

Choose the mitigation after identifying impact. A rollout undo may be safer after a bad release, while a configuration repair or capacity change may be appropriate for another cause. Kubernetes documents both Deployment rollout behavior and running-Pod debugging.

How would you perform a safe Kubernetes rollout?

Define availability and surge limits, a readiness contract, a disruption policy, progressive exposure, service-level stop signals, and a compatible rollback. Include configuration, secrets, schema, and dependency changes; changing only the image is not the whole release.

Practice more production failures in Kubernetes Interview Questions for SRE and Platform Engineers.

Terraform interview questions

Why does Terraform need state?

State maps configuration resource addresses to remote objects and stores information Terraform needs to plan changes. Remote state enables collaboration, but it requires access control, encryption, backup, and locking. Treating state as a disposable cache can lead to duplicate or destructive operations.

Terraform’s state documentation explains these bindings, while state locking protects supported write operations. Never force-unlock until you have verified that no active writer owns the lock.

How do you handle infrastructure drift?

First determine whether the out-of-band change was an emergency fix, an unauthorized edit, or a stale model. Inspect a plan and use refresh-only behavior when you need to reconcile state without proposing remote changes. Then either codify the intended infrastructure or revert the remote object through the normal reviewed path.

Drift detection without ownership creates alert noise. Define who triages it, how emergency changes return to code, and which resources are intentionally ignored. HashiCorp’s resource drift tutorial shows the refresh-only workflow.

How do you make Terraform changes safe at scale?

Separate reusable modules from environment composition, pin versions, validate and plan in CI, restrict apply identity, review destructive changes, lock state, and stage high-risk rollouts by account or region. Avoid one state file whose blast radius spans unrelated systems.

Use the Terraform interview guide for modules, imports, locking, and safe-change scenarios.

Incident and troubleshooting questions

Production errors spike after a deployment. What do you do first?

Use a five-step recovery loop:

  1. Scope impact: affected users, operations, regions, start time, and data or security risk.
  2. Collect evidence: latency, traffic, errors, saturation, recent changes, logs, and traces.
  3. Mitigate reversibly: pause rollout, roll back, shift traffic, disable a feature, or shed load.
  4. Verify the user path: confirm real request success and check delayed work, queues, and reconciliation.
  5. Prevent recurrence: document the timeline and create owned, testable guardrails.

DevOps incident recovery framework from impact and evidence through mitigation verification and prevention

A calm incident answer separates impact reduction from root-cause investigation and verifies recovery from the user’s path.

Google SRE identifies latency, traffic, errors, and saturation as four core monitoring signals. OpenTelemetry distinguishes metrics, logs, and traces; use them as correlated evidence, not three unrelated search boxes.

Tell me about an incident you handled

Structure the story around impact, your role, the decision you made under uncertainty, how you coordinated, and the evidence that proved recovery. Finish with a prevention change and its owner. Avoid a minute-by-minute log dump or claiming sole credit for a team response.

NIST SP 800-61 Rev. 3 frames incident response as part of ongoing risk management, including preparation, detection, response, and recovery. For more scenarios, use the incident response and monitoring and observability guides.

A complete cross-tool scenario

Suppose checkout errors rise five minutes after a release. The pipeline is green, Kubernetes reports the rollout complete, and only one region is affected.

Start by confirming checkout success, error shape, affected region, and change timing. Pause further exposure. Compare application version, configuration, secrets, Terraform changes, cluster events, readiness, dependency health, and resource saturation between healthy and unhealthy regions.

If the new image is the only meaningful difference and rollback is data-compatible, reduce harm by reverting exposure. If Terraform changed a regional dependency, do not rebuild the container to disguise infrastructure drift. If a migration blocks rollback, disable the feature or roll forward with the smallest compatible fix. Finally, verify customer success, drain retry queues, reconcile incomplete transactions, and add a gate that detects the same failure earlier.

This answer is strong because each tool serves a decision. The pipeline proves provenance, Terraform describes infrastructure intent, Kubernetes reports desired and observed workload state, and telemetry tests the user-facing outcome.

Five PracHub questions to practice

The complete question title in the first column opens the prompt and written solution.

PracHub questionPrimary signalFollow-up to add
Design a CI Job Scheduler with Live Build OutputPipeline scheduling, logs, and isolationAdd retries, cancellation, fairness, and artifact retention
Design Safe Configuration and Deployment for Hundreds of ServicesVersioned configuration, staged rollout, and driftAdd emergency changes and regional rollback
Design secure Kubernetes with CI/CDDelivery security and cluster operationsAdd short-lived identity and supply-chain provenance
Present Your Infrastructure and Platform ExperienceTechnical depth and ownershipQuantify impact and name one trade-off you personally made
Describe leading an infrastructure initiativeInfluence, risk, and executionExplain resistance, migration sequencing, and verification

A focused seven-day study plan

  1. Day 1: map one service from commit to production; identify artifacts, identities, gates, and rollback constraints.
  2. Day 2: build and inspect a container; explain layers, runtime state, networking, limits, and persistence.
  3. Day 3: solve three Kubernetes incidents involving readiness, CrashLoopBackOff, and a stalled rollout.
  4. Day 4: review Terraform state, locking, modules, plans, drift, imports, and staged application.
  5. Day 5: diagnose a mixed delivery incident using metrics, logs, traces, events, and recent changes.
  6. Day 6: rehearse two infrastructure stories: one incident and one automation or migration initiative.
  7. Day 7: complete a mixed mock; classify misses as knowledge, evidence, risk, recovery, or communication gaps.

Frequently asked questions

What topics are asked in a DevOps engineer interview?

Common topics include Linux and networking, scripting, Git, CI/CD, containers, Kubernetes, infrastructure as code, cloud permissions, secrets, observability, incident response, and behavioral ownership. The weighting varies by role. Use the current job description and interview invitation to decide which stack deserves the deepest practice.

Do I need to memorize Docker, Kubernetes, and Terraform commands?

Know a small evidence toolkit, but explain what each command would prove and how the result changes your next action. Interviewers usually learn more from a disciplined diagnostic path than from a long command list. Syntax can be checked; unsafe judgment is harder to repair.

How much coding is in a DevOps interview?

It varies. Some roles use data-structure problems; others use scripting, API integration, configuration transformation, concurrency, or debugging. Prepare one fluent language, shell fundamentals, error handling, tests, and complexity. Ask the recruiter whether the exercise is algorithmic, automation-oriented, or tied to a specific environment.

Is a DevOps interview the same as an SRE or platform engineer interview?

No, although the boundaries overlap. DevOps roles often emphasize delivery automation and shared operational ownership; SRE roles emphasize measured reliability and engineering responses to operational work; platform roles emphasize self-service capabilities and developer experience. The actual team charter matters more than the label.

How should I answer when I do not know the exact tool?

State the durable requirement first, describe the mechanism you know, and map it to the unfamiliar tool with explicit assumptions. Ask for the relevant behavior—identity, state, rollout, failure, or observability—instead of bluffing syntax. Strong fundamentals and transparent reasoning are better than fabricated product knowledge.

Final takeaway

The best preparation for DevOps engineer interview questions treats CI/CD, Docker, Kubernetes, Terraform, and incidents as one operating system for change. Use CHANGE to make the artifact, evidence, exposure, health signals, and recovery path visible. Then rehearse failures aloud until you can protect users, test hypotheses, and explain trade-offs without hiding behind tool names.

Sources and Further Reading


Comments (0)