Workload Controllers Tasks — StatefulSet, DaemonSet, Job, CronJob¶
Track A module 6. Do after service-networking-tasks.md.
Read alongside: ../kubernetes/12-workload-controllers.md, ../kubernetes/13-statefulset-deep-dive.md.
Deployment/ReplicaSet you already have. These four cover everything else, and two of them matter directly for GPU work: Jobs are how batch and training workloads run, and DaemonSets are how device plugins and node agents run.
The one idea: every workload controller is the same reconcile loop with a different notion of identity.
Setup: kubectl create ns wl-lab.
Level 0 — Orientation¶
kubectl api-resources --api-group=appsand--api-group=batch- All five own their Pods via ownerReferences — verify with
kubectl get pod <p> -o jsonpath='{.metadata.ownerReferences}'(api-machinery-tasks.mdLevel 3). - Only Deployment has the extra ReplicaSet layer. StatefulSet and DaemonSet own Pods directly, which is why their rollouts behave differently.
Level 1 — StatefulSet¶
- Task 1.1 — Ordinals and the headless Service
apiVersion: v1 kind: Service metadata: {name: db, namespace: wl-lab} spec: clusterIP: None publishNotReadyAddresses: true selector: {app: db} ports: [{port: 5432, name: pg}] --- apiVersion: apps/v1 kind: StatefulSet metadata: {name: db, namespace: wl-lab} spec: serviceName: db replicas: 3 selector: {matchLabels: {app: db}} template: metadata: {labels: {app: db}} spec: containers: - name: c image: busybox command: ["sh","-c","sleep 3600"] volumeMounts: [{name: data, mountPath: /data}] volumeClaimTemplates: - metadata: {name: data} spec: accessModes: [ReadWriteOnce] resources: {requests: {storage: 100Mi}} - Verify:
db-0,db-1,db-2— created in order, each waiting for the previous to be Ready. -
Learn:
serviceNameis required and must point at a headless Service. It's what makes per-Pod DNS work. -
Task 1.2 — Per-Pod DNS
-
Learn:
<pod>.<service>.<ns>.svc.cluster.local. Stable across restarts and reschedules. This is how every clustered database does peer discovery — and whypublishNotReadyAddressesmatters (service-networking-tasks.mdEC-8). -
Task 1.3 — One PVC per Pod, and it survives
- Verify:
hello. Same ordinal → same PVC. -
Learn:
volumeClaimTemplatescreates one PVC per ordinal, named<template>-<sts>-<ordinal>. It binds by name, so identity survives. -
Task 1.4 — Deleting the StatefulSet does NOT delete the PVCs
- Do: delete the StatefulSet, then
kubectl get pvc -n wl-lab. - Verify: all three still there. Recreate the StatefulSet — it reattaches.
-
Learn: deliberate, and a footgun. Data is preserved by default;
persistentVolumeClaimRetentionPolicy(1.27+) lets you opt into deletion. -
Task 1.5 — Rolling update is reverse-ordinal
- Do: change the image, then
kubectl get pods -n wl-lab -w - Verify:
db-2first, thendb-1, thendb-0. -
Learn: highest ordinal first, one at a time, waiting for Ready. For a primary/replica database this updates replicas before the primary — deliberate.
-
Task 1.6 — Partitioned rollout = canary
-
Learn: only ordinals ≥ partition update. Set it to 2, verify the image change, then lower it to roll the rest. This is the built-in staged rollout.
-
Task 1.7 — Parallel pod management
- Learn:
podManagementPolicy: Paralleldrops the ordered guarantee for creation and deletion (not for updates). Right for genuinely peer-to-peer systems where startup order is irrelevant; wrong for anything with a bootstrap sequence.
Level 2 — DaemonSet¶
- Task 2.1 — One per node, automatically
kubectl create -n wl-lab -f - <<'EOF' apiVersion: apps/v1 kind: DaemonSet metadata: {name: agent, namespace: wl-lab} spec: selector: {matchLabels: {app: agent}} template: metadata: {labels: {app: agent}} spec: containers: [{name: c, image: busybox, command: ["sh","-c","sleep 3600"]}] EOF kubectl get ds,pods -n wl-lab -o wide -
Learn: no
replicasfield. The count is derived from the node set. Add a node and a Pod appears; drain one and it goes. -
Task 2.2 — The scheduler still schedules it
- Do:
kubectl get pod -n wl-lab -l app=agent -o jsonpath='{.items[0].spec.affinity}' | jq - Verify: a
nodeAffinitypinning it to one specific node name. -
Learn: modern Kubernetes has the DaemonSet controller create Pods with node affinity and lets the default scheduler place them. It does not bypass scheduling — so a DaemonSet Pod can go Pending for insufficient resources like anything else.
-
Task 2.3 — Reaching tainted nodes
- Do:
kubectl get ds -n kube-system kube-proxy -o jsonpath='{.spec.template.spec.tolerations}' | jq -
Learn: system DaemonSets tolerate nearly everything — including
node.kubernetes.io/not-readyand control-plane taints. A monitoring agent that doesn't tolerate them has blind spots exactly where you need visibility. Seescheduling-constraints-tasks.mdLevel 3. -
Task 2.4 — Targeting a subset
- Do: label a node
accelerator=gpuand addnodeSelector: {accelerator: gpu}. - Verify: Pods only on labelled nodes.
- Learn: this is exactly how the NVIDIA device plugin ships — a DaemonSet
restricted to GPU nodes, advertising
nvidia.com/gputo the kubelet. Yourdevice-plugin-tasks.mdendgame is a DaemonSet.
Level 3 — Job¶
- Task 3.1 — Run to completion
-
Learn: the Pod ends
Completed, notRunning, and is not restarted. Job Pods must userestartPolicy: OnFailureorNever—Alwaysis rejected. -
Task 3.2 — completions and parallelism
- Verify: two at a time until six succeed.
-
Learn:
completions= how many must succeed.parallelism= how many at once. This is the batch primitive underneath most ML training jobs. -
Task 3.3 — Indexed jobs
- Do: read
JOB_COMPLETION_INDEXfrom the env inside a Pod. -
Learn: each Pod gets a stable index 0..N-1 — how you shard work without a queue. This is the shape of distributed training rank assignment.
-
Task 3.4 — backoffLimit and failure
- Do: create a Job with
command: ["sh","-c","exit 1"]andbackoffLimit: 2. - Verify: retries with exponential backoff (10s, 20s, 40s…), then
type: Failed, reason: BackoffLimitExceeded. -
Learn:
backoffLimitcounts Pod failures across the whole Job, not per index. Default 6. -
Task 3.5 — activeDeadlineSeconds and TTL
-
Learn:
activeDeadlineSecondskills the Job regardless of retries — it beatsbackoffLimit.ttlSecondsAfterFinisheddeletes the finished Job and its Pods, and without it, completed Jobs accumulate forever. Those Pods still holdspec.nodeNameand resource requests, which is exactlycontroller-tasks.mdEC-7 — the reason naive capacity dashboards over-report. -
Task 3.6 — Pod failure policy
- Learn: distinguishes "my code is broken, stop retrying" from "the node was preempted, that shouldn't count." Essential on spot/preemptible GPU capacity, where infrastructure churn would otherwise burn your whole backoff budget.
Level 4 — CronJob¶
- Task 4.1 — Schedule
-
Learn: a CronJob creates Jobs, which create Pods. Three levels of ownership.
-
Task 4.2 — concurrencyPolicy
-
Learn:
Allow(default, overlapping runs),Forbid(skip if still running),Replace(kill the old one). A slow job onAllowpiles up until the cluster is full — the classic CronJob outage. -
Task 4.3 — startingDeadlineSeconds and missed runs
-
Learn: if the controller is down past the deadline, runs are skipped. Miss 100 schedules with no deadline set and the CronJob stops permanently with
Cannot determine if job needs to be started. Always set it. -
Task 4.4 — History limits
-
Learn:
successfulJobsHistoryLimit(3) andfailedJobsHistoryLimit(1). This is the CronJob's own garbage collection — separate fromttlSecondsAfterFinishedon the Job. -
Task 4.5 — Timezones
- Learn:
spec.timeZone: "Europe/Yerevan"(1.27+). Without it, schedules use the controller manager's timezone, usually UTC. Every DST bug traces here.
Level 5 — Edge Cases & Production Nuances¶
EC-1 — StatefulSet stuck because Pod 0 won't start¶
- Trap:
db-0is CrashLooping, sodb-1anddb-2are never created. - Why: ordered startup waits for Ready. One broken Pod blocks the whole set.
- Fix: debug
db-0, or switch topodManagementPolicy: Parallelif ordering isn't genuinely required. - Rule: ordering is a guarantee and a serial dependency chain.
EC-2 — StatefulSet PVC keeps stale data¶
- Trap: you delete a broken Pod expecting a clean start; it comes back with the same corrupted volume.
- Why: identity binds Pod ordinal to PVC name. Deleting the Pod changes nothing.
- Fix: delete the PVC too, then the Pod.
- Rule: "delete the pod and see" doesn't work for StatefulSets. That instinct comes from Deployments and it will mislead you here.
EC-3 — DaemonSet Pods Pending forever¶
- Diagnose:
kubectl describe pod→Insufficient cpu, ornode(s) had untolerated taint. - Why: DaemonSets are scheduled normally (Task 2.2). If nodes are already fully booked by requests, the agent doesn't fit.
- Fix: small requests plus a high
priorityClassName(e.g.system-node-critical) so it can preempt. - Rule: node agents must be tiny and high-priority, or they'll be absent from exactly the overloaded nodes you most need to observe.
EC-4 — Completed Job Pods inflate capacity numbers¶
- Trap: a namespace shows high GPU/CPU allocation with nothing running.
- Why: terminal Pods keep
spec.nodeNameandspec.resourcesuntil deleted. They hold no real capacity but appear in naive queries. - Fix:
ttlSecondsAfterFinishedon every Job, and filter onstatus.phasewhen summing. - Rule: the same bug as
controller-tasks.mdEC-7, from the workload side. You'll meet this in real capacity work.
EC-5 — Job retries a broken image forever¶
- Trap:
ImagePullBackOffand thebackoffLimitnever trips. - Why: the Pod never ran, so depending on version it may not count as a Pod failure — it just sits there.
- Fix:
activeDeadlineSecondsas a hard stop. - Rule:
backoffLimitbounds failures; onlyactiveDeadlineSecondsbounds time. Set both.
EC-6 — CronJob stopped silently weeks ago¶
- Diagnose:
kubectl describe cronjob→Cannot determine if job needs to be started: too many missed start times. - Fix: set
startingDeadlineSeconds(e.g. 200) and recreate. - Rule: a CronJob that misses 100 schedules disables itself permanently. Alert
on
status.lastScheduleTimeage — nothing else will tell you.
EC-7 — Two CronJob runs overlap and corrupt state¶
- Trap: a job that usually takes 30s occasionally takes 5 minutes; on a
*/1schedule with defaultAllow, five copies run concurrently. - Fix:
concurrencyPolicy: Forbid. - Rule: the default is the unsafe one. Change it unless overlap is genuinely fine.
Cheat sheet¶
kubectl get sts,ds,job,cronjob -n NS
kubectl rollout status sts/db -n NS
kubectl patch sts db -n NS -p '{"spec":{"updateStrategy":{"rollingUpdate":{"partition":2}}}}'
kubectl get pvc -n NS # data-db-0, data-db-1, ...
dig +short db-0.db.NS.svc.cluster.local # per-pod DNS
kubectl create job NAME --image=IMG -- cmd
kubectl create job manual --from=cronjob/tick -n NS # trigger a CronJob now
kubectl get pods -n NS --field-selector status.phase=Succeeded
kubectl delete pods -n NS --field-selector status.phase==Succeeded
kubectl describe cronjob tick -n NS # missed start times
Mental model to lock in¶
- Identity is the only real difference. None (Deployment), ordinal (StatefulSet), node (DaemonSet), completion (Job), time (CronJob).
- StatefulSet = stable name + stable DNS + stable volume, in creation order, reverse update order. Needs a headless Service.
- StatefulSet PVCs outlive the StatefulSet. Deliberately.
- DaemonSets are scheduled like anything else — they need tolerations, small requests and high priority to be genuinely everywhere.
- Jobs need
ttlSecondsAfterFinished, or terminal Pods pollute capacity views forever. backoffLimitbounds failures,activeDeadlineSecondsbounds time. Set both.- CronJob defaults are unsafe:
Allowconcurrency, nostartingDeadlineSeconds, controller timezone.
CronJob ──schedule──▶ Job ──completions/parallelism──▶ Pods ──▶ Completed
│
└── backoffLimit · activeDeadlineSeconds · ttlSecondsAfterFinished
StatefulSet ──▶ pod-0 ─┐ headless Service ──▶ pod-0.svc, pod-1.svc, ...
pod-1 ─┼── each with data-<sts>-<n> (PVC survives deletion)
pod-2 ─┘ create: 0→N update: N→0
DaemonSet ──▶ one pod per matching node (nodeSelector + tolerations)