Module 6 ยท Lesson 6.2
Requests, Limits & QoS Classes
Every container can declare two numbers per resource (CPU, memory): requests (what it's guaranteed to get) and limits (the ceiling it can't cross). They look similar but do completely different jobs.
Requests: a scheduling decision
The scheduler places a pod on a node only if that node has enough
unreserved capacity to cover the pod's requests โ it never looks at
limits for this decision. Requests are a reservation, exactly like
checking a VM template's minimum RAM before deciding which host has room
for it. Set requests too low and you can over-pack a node; set them too
high and pods sit Pending on a cluster that actually has room.
Limits: a runtime ceiling, enforced differently per resource
Limits are enforced by the kernel's cgroups, and CPU and memory behave differently when you hit them:
- CPU is compressible โ exceed your CPU limit and the kernel just throttles you (slows you down), nothing crashes.
- Memory is not compressible โ there's no "slow down" for memory.
Exceed your memory limit and the Linux kernel's OOM killer kills the
process outright. In a pod, that shows up as
OOMKilled, and the kubelet restarts it per the pod's restart policy โ which can look exactly like a crash loop if the limit is simply set too low for what the app actually needs.
QoS classes: who gets evicted first
Kubernetes derives a Quality of Service class for every pod from its requests/limits, and uses it to decide eviction order when a node runs low on resources:
- BestEffort โ no requests or limits set at all. Evicted first.
- Burstable โ requests set, but not equal to limits (on at least one resource/container). Evicted after BestEffort.
- Guaranteed โ requests equal limits, for every resource, on every container in the pod. Evicted last โ it already reserved exactly what it's allowed to use, so there's no "extra" to reclaim from it cheaply.
Under node memory pressure, QoS class decides who gets evicted first
You don't set QoS class directly โ it's entirely derived from how you set requests/limits, which is exactly what you'll fix in this lesson's lab.
The infrastructure analogy
Requests are like a VM's guaranteed/reserved vCPU and RAM allocation; limits are like a burstable instance's cap. QoS class is the automatic consequence of that choice, similar to how cheaper, oversubscribed instance tiers get reclaimed first when a hypervisor host is under pressure โ except here it's one mechanism, derived automatically from two numbers you were already setting for scheduling reasons.
๐งช Lab: lab-12-resources-and-qos
Preview onlyGoal
See all three QoS classes side by side, and fix one pod so it actually lands in the class it's supposed to be in.
Tasks
- Apply the starting manifests:
kubectl apply -f manifests/ - Check each pod's derived QoS class:
kubectl get pod cache-besteffort api-burstable db-guaranteed -n lab-12-resources-and-qos -o custom-columns=NAME:.metadata.name,QOS:.status.qosClass cache-besteffort(no resources set) andapi-burstable(requests < limits) show the class their names imply.db-guaranteeddoesn't โ it's meant to beGuaranteedbut currently reportsBurstable. Look at itsresourcesblock inmanifests/pods.yamland figure out why.- Fix it: edit
db-guaranteed'sresourcesinmanifests/pods.yamlsorequestsexactly equalslimitsfor bothcpuandmemory. - Resource values on an existing pod are immutable too (same rule as the
container
commandfield in an earlier lab) โ delete it first, then reapply:kubectl delete pod db-guaranteed -n lab-12-resources-and-qos && kubectl apply -f manifests/pods.yaml - Confirm:
db-guaranteednow reports QoS classGuaranteed.
Optional, not graded: try pushing db-guaranteed's memory limit down to
something tiny (e.g. 16Mi) and give it a command that allocates more than
that (sh -c "yes | tr \\\\n x | head -c 100000000 | tail; sleep 3600" is a
crude way to burn memory in busybox). Watch kubectl describe pod report
OOMKilled as the last termination reason โ that's the Linux OOM killer,
not a Kubernetes-level crash.
Check
Run the check once db-guaranteed reports QoS class Guaranteed and the
other two pods are untouched.
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. The scheduler decides where to place a pod based on:
2. A container exceeds its CPU limit vs. exceeds its memory limit. What's the difference?scenario
3. Which requests/limits combination produces 'Guaranteed' QoS?
4. A node is under memory pressure and needs to evict pods. In what order does it evict BestEffort, Burstable, and Guaranteed pods?scenario
5. A pod keeps restarting and `kubectl describe pod` shows 'OOMKilled' as the last termination reason. What's the real fix?scenario
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.