Module 3 ยท Lesson 3.2
Rolling Updates and Rollbacks
How a rolling update actually proceeds
Change the pod template (most often: the image) and the Deployment controller creates a new ReplicaSet, then shifts replicas from old to new according to two settings:
maxUnavailableโ how many pods below desired count you'll tolerate during the rollout (default 25%).maxSurgeโ how many extra pods above desired count you'll allow temporarily (default 25%).
Rolling update, one pod at a time
For 4 replicas, 25% of each rounds to 1 โ so by default you get a strict one-at-a-time replacement: never fewer than 3 ready, briefly up to 5 while a new one comes up alongside the old ones. This is why a rolling update doesn't need a maintenance window: at every point, enough old, known-good pods are still serving traffic.
Revisions and rollback
Every pod-template change gets a revision number, and the old ReplicaSets
stick around (scaled to 0) up to revisionHistoryLimit (default 10) so you
can go back. kubectl rollout undo doesn't "restore a backup" โ it simply
re-points the Deployment's template at an older revision's template, which
the normal rolling-update machinery then rolls out, old-becomes-new, same
as any other update.
Why a bad rollout doesn't usually mean an outage
This is the part worth internalizing: because maxUnavailable keeps most
old pods running, a broken new image (crashing, or in this lesson's lab,
simply unpullable) typically gets stuck, not deployed. Readiness gating
means Kubernetes won't shift more traffic to the new ReplicaSet until the
pods it already created pass their readiness check โ so a bad rollout
stalls partway rather than taking the whole app down. You still have to
notice it's stalled and roll back; it just won't page you for an outage
while you do.
The infrastructure analogy
This is the built-in equivalent of a blue/green or canary deploy script you might have hand-rolled with a load balancer and two target groups โ except the "shift traffic gradually, verify health, stop if it's not healthy" logic is native to the platform, governed by two numbers you set once.
In the lab, you'll run a clean rollout, then a broken one, and practice the exact sequence you'll use in production: notice it's stuck, find out why, roll back.
๐งช Lab: lab-06-rolling-updates
Preview onlyGoal
Run a real rolling update, then run a bad one on purpose, diagnose it while it's stuck, and roll it back.
Tasks
Part 1 โ a clean rolling update
- Apply the starting manifests:
kubectl apply -f manifests/ - Confirm 4/4 pods
Runningonnginx:1.26. - Trigger an update:
kubectl set image deployment/web web=nginx:1.27 -n lab-06-rolling-updates - Watch it roll:
kubectl rollout status deployment/web -n lab-06-rolling-updates kubectl get rs -n lab-06-rolling-updatesโ now there are two ReplicaSets: the old one scaled to 0, the new one at 4. Changing the pod template (the image) is what creates a new revision โ contrast with the scaling-only change in the previous lab.kubectl rollout history deployment/web -n lab-06-rolling-updates
Part 2 โ break it, diagnose it, fix it
- Ship a typo'd tag on purpose:
kubectl set image deployment/web web=nginx:1.27-nope -n lab-06-rolling-updates - Run
kubectl rollout status deployment/web -n lab-06-rolling-updates --timeout=20sโ it will not complete. Note what it says. - Find out why:
kubectl get pods -n lab-06-rolling-updates(look for a pod that isn'tRunning), thenkubectl describe pod <that-pod> -n lab-06-rolling-updatesโ check the Events at the bottom for the real reason. - Because
maxUnavailable: 25%on 4 replicas, the 3 oldnginx:1.27pods are still serving traffic โ a bad rollout doesn't take you down, it just gets stuck. Roll back:kubectl rollout undo deployment/web -n lab-06-rolling-updates - Confirm:
kubectl rollout status deployment/web -n lab-06-rolling-updatescompletes, and every pod is back onnginx:1.27.
Check
Run the check once the rollback has completed.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. With the default maxUnavailable=25% and maxSurge=25% on 4 replicas, how many pods get replaced at a time?
2. What does `kubectl rollout undo` actually do under the hood?
3. You roll out a build with an image tag that doesn't exist. What's the most likely immediate impact on users?scenario
4. Why do old ReplicaSets stick around (scaled to 0) instead of being deleted after a rollout?
5. A rollout has been running for 10 minutes and `kubectl rollout status` hasn't returned. What's the fastest way to find out what's actually wrong?scenario
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.