๐Ÿ“– Read-only preview โ€” lessons, diagrams, and quizzes only. The hands-on labs, live grading, and progress tracking require running the real course locally against a Kubernetes cluster.Run it locally โ†’

Module 6 ยท Lesson 6.1

Liveness, Readiness, Startup Probes

Kubernetes doesn't know if your app is healthy unless you tell it how to check. That's what probes are: small, repeated health checks the kubelet runs against each container, each with a different consequence when it fails.

Three probes, three different jobs

  • Liveness โ€” "is this container still working, or has it wedged?" A failing liveness probe gets the container killed and restarted by the kubelet, right there on the same node. This is self-healing at the container level, one layer below the ReplicaSet-level self-healing you saw in Module 1.
  • Readiness โ€” "should this pod receive traffic right now?" A failing readiness probe does not restart anything โ€” it just pulls the pod out of every Service's endpoints (remember Module 4: Services route only to pods matching their selector and currently Ready). The pod keeps running, Kubernetes just stops sending it requests.
  • Startup โ€” "has this container finished booting yet?" While a startup probe is defined and hasn't yet succeeded, liveness and readiness probes are held off entirely. This exists for slow-starting apps (a JVM with a big warmup, a database replaying a WAL) that would otherwise get killed by liveness for being "unhealthy" when they're really just not done starting.

Readiness vs. liveness: two very different reactions to a failing probe

pod/web-1READY: 1/1RESTARTS: 0
Service โ€œwebโ€ endpoints: [ web-1 ]
Readiness failing only pulls the pod out of Service endpoints โ€” no restart, traffic just stops arriving. Liveness failing gets the container killed and restarted by the kubelet. Startup gates both until the app has had time to boot.

Why readiness failures are the quiet kind

A liveness failure is loud โ€” you'll see RESTARTS climb and kubectl describe pod will show it. A readiness failure is quiet: the pod shows Running, restart count stays at 0, and the only visible symptom is that a Service's endpoint list is shorter than expected (or empty) and traffic silently has nowhere to go. This is one of the most common "it's not working and nothing looks wrong" situations in real clusters โ€” the fix is always the same: check kubectl get endpointslice -l kubernetes.io/service-name=<service> (the modern, non-deprecated equivalent of kubectl get endpoints, per Module 4), then kubectl describe pod on the pods that should be backing it, and look at the probe's own failure message in the Events section.

Getting probe config wrong

All three probes are usually httpGet (hit a path/port and expect 2xx-3xx), tcpSocket (just open a connection), or exec (run a command, 0 means healthy). The most common real-world bug is embarrassingly simple: the probe's path or port doesn't match what the app actually serves โ€” a typo, or a path that changed when the app was updated but the manifest wasn't. That's exactly the bug in this lesson's lab.

The infrastructure analogy

This is the same job a load balancer's health check endpoint does in a traditional setup โ€” except here it's defined per-pod, checked by the kubelet locally rather than by a remote LB polling over the network, and wired automatically into Service routing rather than something you configure separately in two places and hope stays in sync.

๐Ÿงช Lab: lab-11-probes

Preview only

Goal

Diagnose and fix the "quiet" failure mode from the lesson: pods that are Running, never restarting, and still completely unreachable.

Tasks

  1. Apply the starting manifests: kubectl apply -f manifests/
  2. Check pod status: kubectl get pods -n lab-11-probes โ€” notice STATUS is Running but READY is 0/1 for both pods, and restart count is 0.
  3. Confirm the symptom: kubectl get endpoints web -n lab-11-probes โ€” no addresses listed. The Service exists, the pods exist, but nothing is wired together.
  4. Find out why: kubectl describe pod <pod-name> -n lab-11-probes โ€” look at the Events section for a Readiness probe failed (or Unhealthy) message. It names the exact path and status code.
  5. Compare against what nginx actually serves: kubectl exec -n lab-11-probes <pod-name> -- curl -s -o /dev/null -w "%{http_code}\n" http://localhost/healthz vs .../ โ€” one is 404, one is 200.
  6. Fix manifests/deployment.yaml: point the readinessProbe's httpGet.path at a path nginx actually serves (/), then kubectl apply -f manifests/deployment.yaml again.
  7. Confirm: kubectl get pods -n lab-11-probes shows 1/1 Ready for both, and kubectl get endpoints web -n lab-11-probes lists two addresses.

Check

Run the check once both pods are 1/1 Ready and the Service has endpoints.

This lab runs against a real local Kubernetes cluster with an automated grader โ€” clone the repo and run make start to do it for real.

๐Ÿ“ Quiz

1. A container fails its liveness probe. What happens?

2. A pod shows Running, 0 restarts, but a Service that should route to it has no endpoints. What's the most likely cause?scenario

3. What is a startup probe for?

4. You changed an app's health check path from /health to /healthz in the app's code, but forgot to update the Deployment's readinessProbe. What do you see?scenario

5. Which probe type just needs a command to exit 0?

Progress isn't saved in this preview โ€” run the course locally to track completion and grade labs for real.