Skip to content

Kubernetes Deployment Guide for Beginners: From Docker to Production

Hands-on Kubernetes Deployment guide for developers graduating from Docker: manifests, kubectl, minikube, rolling updates, probes, and resource limits.

Comparison card: Kubernetes Deployment Guide For Beginners: From Docker To Production

Kubernetes Deployment is a declarative controller object that manages the lifecycle of containerized application replicas across a cluster, handling rollout, scaling, and self-healing without manual intervention. See also: webhook.

Docker gets a containerized application running on one machine. The problem starts when that machine crashes, traffic spikes past a single instance's capacity, or an update needs to roll out without taking the service offline. Docker alone has no native answer to any of those three situations. Kubernetes addresses all of them through a structured layer of container orchestration that treats your application's desired state as a declaration, not a manual checklist. This guide walks through the local setup, YAML manifest authoring, deployment lifecycle commands, failure diagnosis, and the decision point between staying with Docker and graduating to Kubernetes.

What Kubernetes Does and Why Docker Alone Is Not Enough

Video thumbnail shows Kubernetes Tutorial for Beginners [FULL COURSE in 4 Hours]
Kubernetes Tutorial for Beginners [FULL COURSE in 4 Hours]. Video: TechWorld with Nana via YouTube.

Kubernetes Deployment tracks the desired state of a containerized application and continuously reconciles actual cluster state to match it. Docker, without an orchestrator, offers no automatic restart after a process crash, no mechanism to maintain N simultaneous replicas across hosts, and no native rolling update coordination. Those gaps are not edge cases; they are daily operational realities for any application beyond a single developer's laptop. Container orchestration fills each gap: Kubernetes detects a failed pod and schedules a replacement, distributes replicas across healthy nodes, and executes updates one pod at a time to preserve availability.

Kubernetes splits its workload between two planes. The control plane consists of the API server (the HTTP front end for all cluster operations), etcd (a distributed key-value store that holds cluster state), the scheduler (assigns pods to nodes), and the controller manager (runs reconciliation loops, including the Deployment controller). The data plane consists of Kubernetes nodes running kubelet (the per-node agent) and a container runtime such as containerd or CRI-O. Understanding this split matters before touching <code>kubectl: commands go to the API server; pods execute on nodes.

For more on how Kubernetes fits into a microservices architecture, including how the Kubernetes API server adopts gRPC for internal communication, see the Microservices Communication guide.

Kubernetes Architecture in Plain Terms

A web browser displaying the Kubernetes Dashboard, showing Pods overview with CPU and memory usage graphs.
Credit: Kubernetes

Four terms come up in every Kubernetes conversation. Knowing each one before running a command prevents misreading error messages and kubectl output.

Control Plane
The set of processes (API server, etcd, scheduler, controller manager) that store cluster state and issue scheduling decisions; it does not run application workloads.
Kubernetes Node
A worker machine (virtual or physical) that runs pods; each node runs kubelet, kube-proxy, and a container runtime.
Pod
The smallest deployable unit in Kubernetes; one or more containers sharing a network namespace and storage volumes, always scheduled together on the same node.
Kubelet
The per-node agent that watches the API server for pod assignments and ensures containers are started, healthy, and restarted on failure per the spec.

The Kubernetes Deployment documentation from CNCF contains the authoritative architecture diagram mapping these components to API objects.

Setting Up Your Local Environment

Kubernetes Deployment requires a running cluster to test, and the fastest path to a local one for most developers is minikube. Minikube runs a single-node Kubernetes cluster inside a virtual machine or directly in Docker, which means the container runtime is already present when you start it. The official system requirements are 2 CPU cores and 2 GB of free RAM; below those figures, pod scheduling will stall with Insufficient cpu events.

Docker Desktop ships its own Kubernetes toggle (Settings > Kubernetes > Enable Kubernetes). That toggle is adequate for developers who want the simplest possible setup. A minikube cluster gives more control: you can pin the Kubernetes version, switch container runtimes, and simulate multi-node topologies with minikube node add. For a beginner working through this guide, either option works.

The Docker image build step happens before any cluster interaction, so Docker Engine (or Docker Desktop) is a prerequisite regardless of which local cluster approach you choose. See the Docker Dockerfile reference for build instruction syntax. Once your image exists locally or in a registry, the cluster can pull it.

Installing Docker and kubectl

  1. Install Docker Desktop or Docker Engine for your platform (macOS, Linux, or Windows) from docs.docker.com.
  2. Install kubectl via Homebrew on macOS (brew install kubectl), apt/yum on Linux, or winget on Windows; see kubernetes.io/docs/tasks/tools/.
  3. Install minikube from the minikube Getting Started documentation; the package manager paths match kubectl.
  4. Start the minikube cluster with minikube start --driver=docker; this pulls the Kubernetes node image and configures your kubeconfig automatically.
  5. Verify with kubectl get nodes; a single row showing minikube Ready confirms the cluster is reachable.

Once the minikube cluster reports Ready, kubectl cluster-info prints the control plane URL and CoreDNS addresses. These two commands together confirm both the cluster and your client configuration are working before you write any manifest. For integrating this local workflow into a CI/CD pipeline that builds and pushes Docker images automatically, see CI/CD pipeline and programming languages.

Writing Your First YAML Manifest

Kubernetes Deployment is declared in a YAML manifest: a plain-text file that describes the desired state of the workload. The manifest below is production-aware, not a stripped-down demo. It includes resource limits and a liveness probe because omitting either in a shared cluster causes neighbor-affecting failures, not just local ones.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-app
  labels:
    app: web-app
spec:
  replicas: 3
  selector:
    matchLabels:
      app: web-app
  template:
    metadata:
      labels:
        app: web-app
    spec:
      containers:
        - name: web-app
          image: your-registry/web-app:v1.0.0
          ports:
            - containerPort: 8080
          resources:
            requests:
              cpu: "250m"
              memory: "128Mi"
            limits:
              cpu: "500m"
              memory: "256Mi"
          livenessProbe:
            httpGet:
              path: /healthz
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 10

Each field has a specific function. apiVersion: apps/v1 identifies the API group that owns the Deployment resource. spec.replicas: 3 tells the controller to maintain three pod instances. spec.selector.matchLabels is the label selector that binds the controller to its pods; the label must match spec.template.metadata.labels exactly, because mismatches produce orphaned pods that no controller manages. The deployment configuration's spec.template block is the pod blueprint: every pod the controller creates derives from it.

When you run kubectl apply -f deployment.yaml, the controller creates a replica set, which in turn creates individual pods. The controller owns the replica set; the replica set owns the pods. This two-level ownership enables rollback: each update produces a new replica set, and the prior one stays dormant until explicitly garbage-collected. Running kubectl apply again after editing the manifest is idempotent; Kubernetes computes the diff and applies only what changed.

To see how YAML manifests fit into a broader infrastructure-as-code workflow, including tools that provision the underlying cluster nodes, see infrastructure as code tools.

Resource Limits and Liveness Probes

Resource limits protect shared Kubernetes nodes from runaway containers. The scheduler uses requests values for pod scheduling decisions (bin-packing pods onto nodes with sufficient capacity) and uses limits to cap consumption at runtime. A container without requests.cpu is invisible to the scheduler's bin-packing logic. A container without limits.memory can consume all available memory on a Kubernetes node and trigger OOM kills on unrelated neighboring pods. Both scenarios fail silently until they cause an outage. See the Kubernetes Resource Management documentation for the formal definition of requests vs limits.

A liveness probe tells kubelet whether a container is alive. Without one, kubelet considers any running process healthy, even a deadlocked web server. The initialDelaySeconds field prevents the probe from firing before the application has finished its startup sequence; setting it too low causes CrashLoopBackOff on applications with slow initialization. See the Kubernetes probe configuration documentation for HTTP, TCP, and exec probe types.

FieldPurpose
resources.requests.cpuMinimum CPU guaranteed to the container; used by the scheduler for pod scheduling.
resources.requests.memoryMinimum memory guaranteed; scheduler rejects placement if no node has this free.
resources.limits.cpuCPU cap; container is throttled (not killed) if it exceeds this value.
resources.limits.memoryMemory cap; container is OOM-killed if it exceeds this value.
livenessProbe.httpGetHTTP endpoint kubelet polls to confirm the process is responding.
livenessProbe.initialDelaySecondsSeconds to wait after container start before the first probe fires.

Scale, Rollout, and Rollback with kubectl

Kubernetes Deployment lifecycle operations flow through four kubectl commands. Understanding what each command does to the underlying replica set prevents confusion when reading rollout history or diagnosing failed updates.

Scaling and Update Strategies

  1. kubectl apply -f deployment.yaml: applies the deployment configuration from the manifest; creates or updates the object; triggers a rolling update if the pod template changed. Output: deployment.apps/web-app configured.
  2. kubectl rollout status deployment/web-app: streams the rolling update progress until all pods in the new replica set are Ready; exits with code 0 on success, code 1 on timeout. Output: deployment "web-app" successfully rolled out.
  3. kubectl scale deployment/web-app --replicas=5: imperatively adjusts the replica count without touching the manifest; useful for traffic spikes. Output: deployment.apps/web-app scaled.
  4. kubectl rollout undo deployment/web-app: reverts to the previous replica set; Kubernetes keeps the prior ReplicaSet dormant for exactly this purpose; the revisionHistoryLimit field (default: 10) controls how many prior ReplicaSets are retained.

A rolling update by default keeps at least 75% of desired pods available (maxUnavailable: 25%) and allows at most 125% of desired pods to exist simultaneously (maxSurge: 25%). Override both in spec.strategy.rollingUpdate when a tighter availability guarantee or a faster cutover is required. For infrastructure provisioning context, including the Terraform workflows that create the underlying node pools on EKS or GKE, see Terraform vs Ansible vs CloudFormation.

StrategyDowntimeUse Case
RollingUpdate (default)None (with correct maxUnavailable/maxSurge)Stateless services with multiple replicas; standard production deployments.
RecreateFull gap between old and new podsSingle-replica stateful applications where two versions cannot run simultaneously.

Manual kubectl apply is a development workflow. Production pipelines push through GitHub Actions or GitLab CI, where the same command runs as a pipeline step after image build and test. The rollout and rollback commands are identical in both contexts; only the trigger changes. See also: full stack development.

Troubleshooting Common Kubernetes Issues

Kubernetes Deployment surfaces failures through pod status strings and event logs. Six failure states account for the majority of beginner problems; each has a specific diagnostic command and a narrow set of root causes.

ImagePullBackOff
Diagnostic: kubectl describe pod <pod-name>: look for "Failed to pull image" in Events.
Root cause: Wrong image tag, or private registry credentials not mounted as an imagePullSecret.
Fix: Correct the tag in the deployment configuration, or create a Secret with registry credentials and reference it in spec.imagePullSecrets.
CrashLoopBackOff
Diagnostic: kubectl logs <pod-name> --previous shows the last exit output before the crash.
Root cause: Application exits on startup (misconfigured environment variable, missing dependency) or a liveness probe fires before the app is ready because initialDelaySeconds is too short.
Fix: Fix the application error or increase livenessProbe.initialDelaySeconds. See the Kubernetes probe configuration documentation for probe timing guidance.
Pending pods
Diagnostic: kubectl describe pod <pod-name>: Events section shows "Insufficient cpu" or "Insufficient memory" when pod scheduling is blocked.
Root cause: No Kubernetes node has enough unallocated CPU or memory to satisfy the pod's resource limits requests; or a node taint rejects the pod.
Fix: Reduce resource limits requests, add a Kubernetes node to the cluster, or add a matching toleration for the taint.
OOMKilled
Diagnostic: kubectl describe pod <pod-name> shows OOMKilled as the last state reason.
Root cause: Container consumed more memory than its resource limits ceiling; the kernel OOM killer terminated the process.
Fix: Raise limits.memory or fix the memory leak in the application code.
Service not reachable
Diagnostic: kubectl get endpoints <service-name>: an empty Endpoints object (no pod IPs listed) confirms a label selector mismatch.
Root cause: The Service's spec.selector labels do not match the pod labels set by the deployment configuration's pod template.
Fix: Align the label keys and values in both the Service selector and the pod template.
ErrImageNeverPull
Diagnostic: kubectl describe pod <pod-name> shows "ErrImageNeverPull" in Events.
Root cause: imagePullPolicy: Never is set, but the Docker image is not cached on the target Kubernetes node, which is common when switching from a minikube cluster to a remote cluster.
Fix: Push the Docker image to a registry and change imagePullPolicy to Always or IfNotPresent.

Docker vs Kubernetes: Choosing the Right Tool

Kubernetes Deployment adds meaningful operational capability over Docker alone, and with it comes meaningful operational overhead. For teams evaluating container orchestration for the first time, the decision comes down to five concrete requirements.

CapabilityDocker alone (or Compose)With Kubernetes
Replica managementManual restart; single-host onlyAutomatic ReplicaSet maintenance across nodes
Rolling updatesNo native support; requires custom scriptsDeclarative strategy with configurable maxUnavailable and maxSurge
Self-healingNone; crashed containers stay down until manually restartedLiveness probe failures trigger automatic container restarts
Operational overheadLow; single host, no control planeMedium to high; cluster management, etcd backups, RBAC, networking
Scale ceilingOne host (vertical scaling only)Horizontal scaling across multiple Kubernetes nodes

The practical decision rule: use Docker Compose for single-host development, local testing, and CI runners where the containerized application never needs more than one instance. Graduate to Kubernetes when any of the following becomes true: the service needs more than one replica for availability, rolling updates with zero downtime are a requirement, or the workload must run across more than one physical or virtual host.

A minikube cluster covers the entire learning curve from manifest authoring through rollback without touching cloud infrastructure. Once local workflows are stable, the natural next step is a managed cluster. For a comparison of AWS deployment options including EKS (Elastic Kubernetes Service), see AWS deployment options including EKS.

Further reading

Share this guide

Marcus Vetri

Marcus Vetri covers developer tools and enterprise software for techshooked: the IDEs, package managers, build systems, and runtimes that engineers keep open all day. He writes comparison-first and reproducibility-first, stating the version tested, showing the configuration, and separating a real workflow improvement from a marketing claim.