Module 13 ยท Lesson 13.1
Logs, Events, Describe, Metrics
You've already been using these tools since Module 2, one at a time, as each lab needed them. This lesson is the explicit reference: what each one is actually for, so you reach for the right one first instead of guessing.
The four tools
kubectl logs <pod>โ stdout/stderr from a container. Add-c <name>for a specific container in a multi-container pod, and--previousto read the last container's logs after it's already been restarted โ without--previous, you're reading the new container's (probably empty) logs, not the one that crashed.kubectl describe <kind> <name>โ the object's full status plus a chronological Events list at the bottom. For pods specifically, the Last State block (with a Reason likeOOMKilled,Error, orCompleted, and an Exit Code) is often the single fastest way to know what actually happened.kubectl get events --sort-by=.lastTimestampโ the same eventsdescribeshows for one object, but cluster- or namespace-wide and in order. Useful when you don't yet know which object to describe.kubectl top nodes/kubectl top podsโ live CPU/memory usage, sourced from the metrics-server add-on running in this cluster (not part of the Kubernetes API server itself โ iftopever errors out, that's a metrics-server problem, not a sign anything else is broken).
A new failure mode: OOMKilled
This lesson's lab introduces one you haven't hit yet: OOMKilled. Every
container has a memory limit (Module 6); exceed it even briefly and the
Linux kernel's out-of-memory killer ends the process immediately โ there's
no graceful warning, no SIGTERM first, just gone, then restarted by the
kubelet per the pod's restart policy. kubectl describe pod shows this
unambiguously: Last State: Terminated, Reason: OOMKilled. This is
different from a crash (the app itself exiting) or a failed probe (the app
is fine, kubelet thinks otherwise) โ the kernel intervened because of a
hard resource ceiling, which is why the fix lives in the resources block,
not the code.
Reading the chain, not just the symptom
None of these four tools replace understanding what you're looking at.
describe's Events are chronological but not always causally obvious โ
"Scheduled" then "Pulled" then "Created" then "Started" then, seconds
later, "Killing" tells a story, but you have to read it as one. The next
lesson turns this into an explicit method instead of a vague instinct.
In the lab, a Deployment is unhealthy and restarting. You'll use exactly these four tools โ no new YAML, just investigation โ to find out why.
๐งช Lab: lab-26-logs-events-metrics
Preview onlyGoal
A Deployment is unhealthy and keeps restarting. Using only the diagnostic toolkit from this lesson (no new YAML concepts), find out exactly why, then fix it.
Tasks
- Apply the starting manifests:
kubectl apply -f manifests/ kubectl get pods -n lab-26-logs-events-metrics -wโ watch it restart a few times. Note theRESTARTScount climbing.kubectl describe pod -n lab-26-logs-events-metrics -l app=memhogโ scroll to Last State. What's the Reason? What's the Exit Code?kubectl get events -n lab-26-logs-events-metrics --sort-by=.lastTimestampโ find the event that corresponds to what you saw in step 3.- (Optional, timing-dependent)
kubectl top pod -n lab-26-logs-events-metricsโ you may or may not catch it mid-allocation before it's killed again; either way isn't the deciding evidence here,describealready gave you the answer. - Open
manifests/deployment.yaml. The container allocates roughly 200Mi of memory on startup (see itscommand). Fix the mismatch between what it needs and whatresources.limits.memoryallows, thenkubectl apply -f manifests/deployment.yamlagain.
Check
Run the check once the pod has been stably Running (not currently
restarting) for a few seconds.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. A pod just restarted. You want to see the output from the container that crashed, not the fresh one that replaced it. What do you run?
2. kubectl describe pod shows Last State: Terminated, Reason: OOMKilled. What actually happened?scenario
3. Where does kubectl top get its numbers from?
4. You don't yet know which object is causing a problem. What's the best first command?
5. A multi-container pod (app + a logging sidecar) is unhealthy. kubectl logs <pod> shows nothing useful. What should you check next?scenario
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.