๐Ÿ“– Read-only preview โ€” lessons, diagrams, and quizzes only. The hands-on labs, live grading, and progress tracking require running the real course locally against a Kubernetes cluster.Run it locally โ†’

Module 14 ยท Lesson 14.1

Horizontal Pod Autoscaler

Scaling by count, automatically

A HorizontalPodAutoscaler (HPA) is a controller โ€” same reconciliation pattern as everything else in this course โ€” that watches a metric and adjusts a Deployment's (or StatefulSet's) replica count to keep that metric near a target. The most common target: CPU utilization as a percentage of each pod's requested CPU.

That last word matters more than it looks like it should: HPA computes percentage of request, not an absolute number. A pod with no CPU request has nothing to divide by โ€” HPA can't compute a percentage of zero, so it simply can't scale that workload on CPU at all. This is why Module 6's "always set requests" advice wasn't just about scheduling and QoS; it's a hard prerequisite for autoscaling too.

Deployment โ†’ ReplicaSet โ†’ Pods

Deployment: web

spec.replicas = 3

scale

ReplicaSet: web-7d9f6c

desired: 3 ยท current: 3

๐Ÿ“ฆ podweb-1
๐Ÿ“ฆ podweb-2
๐Ÿ“ฆ podweb-3
Scale it and watch: you only ever edit the Deployment's replica count. The ReplicaSet controller is what notices the gap and creates or deletes pods to close it.

The loop, concretely

  1. metrics-server scrapes each kubelet's cAdvisor stats and exposes them through the metrics.k8s.io API โ€” this is a separate, optional add-on, not something the API server does natively.
  2. The HPA controller polls that API on a sync period (default ~15s), averages utilization across the target's pods.
  3. It computes desired replicas โ‰ˆ current ร— (currentMetric / targetMetric), clamped between minReplicas and maxReplicas.
  4. It writes a new replicas value onto the Deployment โ€” which the ReplicaSet controller from Module 3 then reconciles exactly like a manual scale, because to that controller, it is a manual scale. HPA doesn't talk to pods directly; it only ever edits the one number.

The API is autoscaling/v2 โ€” the older v1 (CPU-only, no behavior tuning) and v2beta1/v2beta2 are gone from current Kubernetes; v2 has supported CPU, memory, and custom/external metrics, plus configurable scale-up/scale-down behavior, for a long time now.

Why this feels slow the first time you watch it

Nothing is instant here: the kubelet scrapes periodically, metrics-server scrapes the kubelet periodically, the HPA polls metrics-server periodically, and a new pod still needs to schedule, pull its image, and pass readiness before it counts. Budget a minute or two for a real scale-up, not seconds โ€” that's normal, not a bug, and it's exactly what you'll see in this lesson's lab.

The infrastructure analogy

This is the direct platform equivalent of a cloud provider's Auto Scaling Group scaling policy on a CloudWatch metric โ€” same shape (observe a metric, compare to target, adjust capacity), just one layer down: scaling pod count inside a node pool instead of scaling instance count inside a VPC. Module 15 covers the layer above this โ€” scaling the nodes themselves.

๐Ÿงช Lab: lab-28-hpa

Preview only

Goal

Watch a real HorizontalPodAutoscaler scale a real workload, and fix the one thing most first HPAs get wrong: forgetting the CPU request HPA needs to compute a percentage against.

Tasks

  1. Apply the starting manifests: kubectl apply -f manifests/
  2. Watch the HPA: kubectl get hpa -n lab-28-hpa -w โ€” the TARGETS column will show <unknown>/50% instead of a real percentage.
  3. Run kubectl describe hpa burner -n lab-28-hpa and read the Events โ€” it names the exact problem.
  4. The burner container already busy-loops burning CPU on purpose (see the comment in manifests/deployment.yaml) โ€” there's load, HPA just can't measure it as a percentage of nothing. Add a CPU request (and, good practice, a limit) to manifests/deployment.yaml, then kubectl apply -f manifests/deployment.yaml again.
  5. Watch kubectl get hpa -n lab-28-hpa -w again โ€” within a minute or two TARGETS should show a real, high percentage (the busy-loop blows way past 50% almost immediately), and kubectl get pods -n lab-28-hpa should show more than 1 burner pod.

Check

Run the check once you see more than 1 burner pod. Give it time โ€” this check genuinely waits for a real scale-up (up to a few minutes), it's not instant.

This lab runs against a real local Kubernetes cluster with an automated grader โ€” clone the repo and run make start to do it for real.

๐Ÿ“ Quiz

1. You create an HPA targeting 50% CPU utilization for a Deployment whose containers have no `resources.requests.cpu` set. What happens?scenario

2. What does metrics-server actually do?

3. When an HPA decides to scale up, what does it actually change?

4. Which HPA API version should you use today?

5. You trigger load against an HPA-managed Deployment and immediately check replica count โ€” it's unchanged. Is something broken?scenario

Progress isn't saved in this preview โ€” run the course locally to track completion and grade labs for real.