Module 14 ยท Lesson 14.1
Horizontal Pod Autoscaler
Scaling by count, automatically
A HorizontalPodAutoscaler (HPA) is a controller โ same reconciliation pattern as everything else in this course โ that watches a metric and adjusts a Deployment's (or StatefulSet's) replica count to keep that metric near a target. The most common target: CPU utilization as a percentage of each pod's requested CPU.
That last word matters more than it looks like it should: HPA computes percentage of request, not an absolute number. A pod with no CPU request has nothing to divide by โ HPA can't compute a percentage of zero, so it simply can't scale that workload on CPU at all. This is why Module 6's "always set requests" advice wasn't just about scheduling and QoS; it's a hard prerequisite for autoscaling too.
Deployment โ ReplicaSet โ Pods
Deployment: web
spec.replicas = 3
ReplicaSet: web-7d9f6c
desired: 3 ยท current: 3
The loop, concretely
- metrics-server scrapes each kubelet's cAdvisor stats and exposes
them through the
metrics.k8s.ioAPI โ this is a separate, optional add-on, not something the API server does natively. - The HPA controller polls that API on a sync period (default ~15s), averages utilization across the target's pods.
- It computes desired replicas โ
current ร (currentMetric / targetMetric), clamped betweenminReplicasandmaxReplicas. - It writes a new
replicasvalue onto the Deployment โ which the ReplicaSet controller from Module 3 then reconciles exactly like a manual scale, because to that controller, it is a manual scale. HPA doesn't talk to pods directly; it only ever edits the one number.
The API is autoscaling/v2 โ the older v1 (CPU-only, no behavior
tuning) and v2beta1/v2beta2 are gone from current Kubernetes; v2
has supported CPU, memory, and custom/external metrics, plus configurable
scale-up/scale-down behavior, for a long time now.
Why this feels slow the first time you watch it
Nothing is instant here: the kubelet scrapes periodically, metrics-server scrapes the kubelet periodically, the HPA polls metrics-server periodically, and a new pod still needs to schedule, pull its image, and pass readiness before it counts. Budget a minute or two for a real scale-up, not seconds โ that's normal, not a bug, and it's exactly what you'll see in this lesson's lab.
The infrastructure analogy
This is the direct platform equivalent of a cloud provider's Auto Scaling Group scaling policy on a CloudWatch metric โ same shape (observe a metric, compare to target, adjust capacity), just one layer down: scaling pod count inside a node pool instead of scaling instance count inside a VPC. Module 15 covers the layer above this โ scaling the nodes themselves.
๐งช Lab: lab-28-hpa
Preview onlyGoal
Watch a real HorizontalPodAutoscaler scale a real workload, and fix the one thing most first HPAs get wrong: forgetting the CPU request HPA needs to compute a percentage against.
Tasks
- Apply the starting manifests:
kubectl apply -f manifests/ - Watch the HPA:
kubectl get hpa -n lab-28-hpa -wโ theTARGETScolumn will show<unknown>/50%instead of a real percentage. - Run
kubectl describe hpa burner -n lab-28-hpaand read the Events โ it names the exact problem. - The
burnercontainer already busy-loops burning CPU on purpose (see the comment inmanifests/deployment.yaml) โ there's load, HPA just can't measure it as a percentage of nothing. Add a CPU request (and, good practice, a limit) tomanifests/deployment.yaml, thenkubectl apply -f manifests/deployment.yamlagain. - Watch
kubectl get hpa -n lab-28-hpa -wagain โ within a minute or twoTARGETSshould show a real, high percentage (the busy-loop blows way past 50% almost immediately), andkubectl get pods -n lab-28-hpashould show more than 1burnerpod.
Check
Run the check once you see more than 1 burner pod. Give it time โ this
check genuinely waits for a real scale-up (up to a few minutes), it's not
instant.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. You create an HPA targeting 50% CPU utilization for a Deployment whose containers have no `resources.requests.cpu` set. What happens?scenario
2. What does metrics-server actually do?
3. When an HPA decides to scale up, what does it actually change?
4. Which HPA API version should you use today?
5. You trigger load against an HPA-managed Deployment and immediately check replica count โ it's unchanged. Is something broken?scenario
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.