Module 9 ยท Lesson 9.1
Jobs & CronJobs
Jobs: run-to-completion, not run-forever
Every workload you've built so far โ Deployments, the ReplicaSets behind them โ assumes pods run forever. If one exits, that's treated as a problem, and it gets restarted. A Job flips that assumption: its pods are supposed to finish, and a successful exit is the goal, not a failure.
A Job tracks completions (how many successful pod runs it needs) and
parallelism (how many can run at once). If a pod's container exits
non-zero, the Job retries a new pod, up to backoffLimit โ cross that, and
the Job gives up and marks itself Failed, visible in kubectl get jobs
and kubectl describe job.
This is the same exit-code/Events diagnosis muscle from Module 2's CrashLoopBackOff lab โ the difference is just what "the pod exited" means for the controller watching it: a Deployment sees it as drift to correct forever; a Job sees it as one attempt, counted against a limit.
CronJobs: a Job template on a schedule
A CronJob doesn't run anything itself โ it creates a new Job from its
template on a schedule, using standard cron syntax (minute hour day-of-month month day-of-week, e.g. */5 * * * * for every 5 minutes).
Linux aside: cron is the traditional Unix/Linux job scheduler โ a daemon that reads schedule files and runs commands at the times they specify. Kubernetes didn't invent this syntax; it deliberately reused the one generations of Linux admins already know, instead of designing a new scheduling language.
Each firing creates a fresh Job (and fresh pods) โ nothing persists between
runs unless you explicitly give it somewhere to write (a volume, an
external store). concurrencyPolicy controls what happens if a run is
still going when the next one is due: Allow (both run), Forbid (skip
the new one), or Replace (kill the old one first).
The infrastructure analogy
A Job is your traditional one-off batch script or CI job run โ something with a defined "done." A CronJob is literally what it sounds like: cron, but the "box" it runs on is created fresh for every single run instead of being a long-lived machine you SSH into and schedule things on. No drift between runs, because there's no persistent machine to drift.
In the lab, you'll diagnose a Job that keeps failing, then verify a CronJob you create actually fires on schedule โ not just that you wrote the YAML correctly, but that the cluster really ran it.
๐งช Lab: lab-18-jobs-and-cronjobs
Preview onlyGoal
Diagnose and fix a failing Job using the same exit-code/Events muscle from Module 2, then prove a CronJob actually fires on schedule โ not just that the YAML looks right.
Tasks
Part 1 โ fix the failing Job
- Apply the namespace and the Job:
kubectl apply -f manifests/00-namespace.yaml -f manifests/job.yaml - Watch it:
kubectl get job report -n lab-18-jobs-and-cronjobs -w(Ctrl-C once it settles) - It gives up and shows
FAILED. Find out why:kubectl describe job report -n lab-18-jobs-and-cronjobs(check the Events) andkubectl logs -n lab-18-jobs-and-cronjobs -l job-name=report --tail=20. - The container's command is the bug (see the comment in
manifests/job.yaml). Jobs are immutable once created โ fix the command in the file, then:kubectl delete job report -n lab-18-jobs-and-cronjobs && kubectl apply -f manifests/job.yaml - Confirm:
kubectl get job report -n lab-18-jobs-and-cronjobsshowsCOMPLETIONS 1/1.
Part 2 โ prove the CronJob actually runs
- Apply it:
kubectl apply -f manifests/cronjob.yaml - Wait about 70 seconds (it's scheduled every minute), then check:
kubectl get cronjob heartbeat -n lab-18-jobs-and-cronjobs - Look at
LAST SCHEDULEandkubectl get jobs -n lab-18-jobs-and-cronjobsโ you should see at least one Job the CronJob created on its own, with nokubectl applyfrom you triggering it.
Check
Run the check once the Job is complete and the CronJob has fired at least once (give it the full ~70 seconds โ don't run the check immediately after applying the CronJob).
This lab runs against a real local Kubernetes cluster with an automated grader โ clone the repo and run make start to do it for real.
๐ Quiz
1. A pod managed by a Deployment exits with code 0. A pod managed by a Job exits with code 0. How does each controller react?
2. A Job's pod keeps exiting with a nonzero code. After several retries, `kubectl get jobs` shows it as Failed. What determined how many retries happened before giving up?scenario
3. What does a CronJob actually do when its schedule fires?
4. A CronJob's previous run is still executing when the next scheduled time arrives. With `concurrencyPolicy: Forbid`, what happens?scenario
5. Why does Kubernetes use traditional cron syntax (`*/5 * * * *`) for CronJob schedules instead of something Kubernetes-specific?
Progress isn't saved in this preview โ run the course locally to track completion and grade labs for real.