Module 13 ยท Lesson 13.2
A Structured Debugging Method
You've now independently diagnosed a CrashLoopBackOff (Module 2), a zero-endpoint Service (Module 4), a failing probe (Module 6), a stuck PVC (Module 7), a Forbidden RBAC error (Module 11), and an OOMKill (this module). That's not six unrelated tricks โ it's one method, applied six times. Let's name it.
The method
- Identify the symptom, at the object showing it. Don't start by
reading code or guessing. Start with
kubectl get events/describeon the object that's visibly unhealthy โ it almost always tells you the symptom directly (CrashLoopBackOff,Pending,0 endpoints,Forbidden). - Walk the ownership/dependency chain if the symptom isn't the cause. A Service with no endpoints isn't broken in the Service โ walk down to the Pods it's supposed to select. A Deployment stuck mid-rollout isn't broken in the Deployment โ walk down to the ReplicaSet, then the Pods. You're retracing the exact ownerReference chains from Module 3.
- Name the specific layer that's actually broken. By this point you have a short, known list of candidates: scheduling (taints, resources, affinity), image (tag, pull access), config (ConfigMap/Secret reference), probe, resource limit, RBAC, network/DNS, storage. Pick one based on the evidence, not a hunch.
- Form the smallest possible fix, apply it, re-check. Don't batch
five guesses into one
apply. Change one thing, re-rundescribeorget events, confirm the specific symptom from step 1 is gone before moving on.
A structured debugging method
Why step order matters
Steps 1 and 2 are investigation; step 3 is diagnosis; step 4 is treatment. Skipping straight to step 4 โ "let's just try restarting it" โ sometimes works by accident and teaches you nothing for next time. Worse, on a multi-fault problem (this lesson's lab has more than one independent bug at once) a lucky restart can mask one fault while another is still live, and you'll find out later, under worse conditions.
The infrastructure analogy
This is the same discipline as any good incident-response runbook: symptom first, blast radius / dependency chain next, root cause named explicitly, then one change at a time โ never "restart everything and see." The only Kubernetes-specific part is where to look at each step; the discipline itself is the same one you already use (or should) on traditional infra.
The lab has three independent bugs hiding in one small set of objects. Find and fix all three, using the method above โ not by poking randomly until things turn green.
๐งช Lab: lab-27-structured-debugging
Preview onlyGoal
Three independent bugs are hiding in this namespace. Find and fix all three using the method from this lesson โ symptom, chain, layer, smallest fix โ not by randomly editing things until the check passes.
Tasks
- Apply the starting manifests:
kubectl apply -f manifests/ - Survey:
kubectl get all -n lab-27-structured-debugging. Nothing here is fully healthy. Don't fix anything yet โ just note what looks wrong (a Deployment not at desired replicas, a Service, a Deployment that's Running but not Ready). - Bug 1 โ
webDeployment:kubectl describe pod -n lab-27-structured-debugging -l app=web. What's the container state and reason? Trace it tomanifests/app.yaml's ConfigMap reference and fix the mismatch. - Bug 2 โ
web-svcService:kubectl get endpoints web-svc -n lab-27-structured-debugging. Compare the Service'sselectoragainst thewebDeployment's actual pod labels (kubectl get pods -n lab-27-structured-debugging --show-labels). Fix whichever one doesn't match the other. - Bug 3 โ
workerDeployment: it'sRunningbutREADY 0/1. That's a probe, not a crash โkubectl describe pod -n lab-27-structured-debugging -l app=workerand look at the readiness probe's failure message, then check what path the container's command actually creates vs. what the probe checks. - Reapply whichever files you changed, then re-run
kubectl get all -n lab-27-structured-debuggingand confirm all three are healthy.
Check
Run the check once you believe all three are fixed โ it grades each bug independently.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. A Service has zero endpoints. Where should step 1 (identify the symptom) point you?
2. You suspect three possible causes for a stuck rollout. What does step 4 of the method say to do?scenario
3. Why does step 3 (name the specific layer) come before step 4 (apply a fix)?
4. You have a Deployment with a failing readiness probe AND a separate Service with a selector typo, both in the same broken demo. You fix only the probe. What do you expect?scenario
5. What's the traditional-infra equivalent of this four-step method?
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.