Module 1 ยท Lesson 1.2
The Reconciliation Loop
This is the single most important idea in Kubernetes. Almost everything else โ Deployments, Services, autoscalers โ is a specific instance of this one pattern.
Desired state vs. actual state
When you kubectl apply a manifest, you're not running a command that
does something once. You're writing a row into etcd that says "this is
what I want to exist." That's the desired state.
Separately, the cluster has an actual state: what's really running right now, as reported by kubelets.
A controller is a small, dumb, relentless loop:
for {
observed := get_actual_state()
desired := get_desired_state()
if observed != desired {
take_action_to_close_the_gap()
}
}
The reconciliation loop
There's a Deployment controller, a ReplicaSet controller, a Node controller, a Job controller โ dozens of them, each watching one slice of the API and running this same loop, forever. None of them ever declare victory and stop.
Why this matters for troubleshooting
This reframes almost every production issue you'll debug: something isn't
"broken," it's that actual state can't reach desired state, and your job
is to find out why. A Pod stuck in Pending? The scheduler's loop can't
find a node satisfying the desired constraints. A Deployment stuck
mid-rollout? The ReplicaSet controller's loop is blocked on a readiness
check. You'll come back to this framing constantly, especially in the
observability module.
The traditional-infra contrast
Think about a config-management run (Ansible/Puppet/Chef): you execute it, it converges the box, then it stops until you run it again โ drift can creep back in between runs. Kubernetes controllers never stop running. If something drifts โ a pod gets OOM-killed, a node dies โ the gap is noticed and closed automatically, usually within seconds, with no cron job or human involved.
In the lab, you'll kill something the control plane manages and watch it come back without you lifting a finger a second time.
๐งช Lab: lab-02-reconciliation
Preview onlyGoal
Prove to yourself that the reconciliation loop is real, not just a diagram โ by deliberately breaking something the control plane manages, and watching it fix itself without you touching it again.
CoreDNS (cluster DNS) runs as a Deployment in kube-system. It's a
controller-managed workload exactly like any app Deployment you'll create
later. You're going to kill one of its pods directly.
Tasks
kubectl get pods -n kube-system -l k8s-app=kube-dnsโ note the exact name of one CoreDNS pod.- Open a second terminal and run
kubectl get pods -n kube-system -l k8s-app=kube-dns -wto watch it live. - Back in the first terminal, delete that pod:
kubectl delete pod -n kube-system <pod-name>. - Watch the second terminal. You did not create a replacement pod โ the ReplicaSet controller behind the CoreDNS Deployment noticed the gap between desired (2 replicas) and actual (1) and closed it itself.
- Run
kubectl get events -n kube-system --sort-by=.lastTimestamp | tail -n 15and find theKillingevent for the pod you deleted, followed by aSuccessfulCreatefor its replacement.
Check
Run the check after the replacement pod is Running.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. When you run `kubectl apply -f deploy.yaml`, what actually happens?
2. A controller's reconciliation loop is best described as:
3. You manually edit a Deployment-managed pod's container image with `kubectl edit pod <name>` directly (not the Deployment). What happens shortly after?scenario
4. Compared to running Ansible/Puppet/Chef by hand or on a schedule, Kubernetes's reconciliation loop is different because:
5. A Pod has been stuck in `Pending` for 10 minutes. Through the reconciliation-loop lens, what's the most useful way to frame this?scenario
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.