Skip to content

Resource Management and QoS

A Kubernetes node is a shared multi-tenant Linux machine. The kernel doesn't know what a Pod is — it knows processes, cgroups, page-cache pages, NUMA banks, file descriptors, and PIDs. Every "resource limit" in a Pod spec is a knob on a cgroup, an oom_score_adj byte in /proc, a cpuset mask, or a kubelet decision to evict before the kernel kills the wrong process. Get any of those wrong and your workload silently loses 30% of its throughput, or a noisy neighbor takes the node down at 03:00 with a cascading OOM.

This chapter is about the policies the kubelet enforces and the kernel mechanisms they ride on. We start with the fundamental dichotomy — requests drive scheduling, limits drive enforcement (§2) — then go through CPU semantics (§4) including CFS throttling (§9) and the strong case for not setting CPU limits (§10), memory semantics (§5) including memory.high (§7) and the MemoryQoS feature gate, the three QoS classes (§3) and their oom_score_adj (§8), the static CPU manager and memory manager (§11–12), the topology manager (§13), namespace-level governance via ResourceQuota, LimitRange, and PriorityClass (§14–17), extended resources and hugepages (§18–19), ephemeral storage and PID pressure (§20–21), the kubelet's reservation model and the node cgroup hierarchy (§22–23), eviction signals and ranking (§24–26), in-place pod resize (§28), the right-sizing workflow (§30), observability (§31), and the 25+ ways this all goes wrong in production (§32).

Pre-reqs: chapter 00 (Linux primitives — cgroup v2, OOM killer, CFS scheduler), chapter 10 (kubelet internals — eviction manager, CPU/memory/topology managers), chapter 11 (Pod internals — the object whose .spec.resources we are interpreting). Forward-references: chapter 22 (autoscaling — HPA/VPA/Karpenter consume the same requests/limits), chapter 25 (multi-tenancy — ResourceQuota is the cap that makes shared clusters survivable), chapter 35 (perf — overcommit math at 5000-node scale).


Table of Contents

  1. Why Requests and Limits Exist
  2. The Two-Knob Model: Scheduling vs Enforcement
  3. QoS Class Derivation 3A. CPU Fundamentals: Logical vs Physical, CPU-Seconds, CFS, and Concurrency 3B. Capacity Metrics: Usage, Utilization, Core-Hours, and Memory Accounting
  4. CPU Semantics: Fractional Requests, cpu.weight, cpu.max
  5. Memory Semantics: Bytes, Reservation, memory.max
  6. The cgroup-v2 Tree Under kubepods.slice
  7. memory.high and the MemoryQoS Feature Gate
  8. OOM Scoring: Per-QoS oom_score_adj
  9. CFS Quota and CPU Throttling
  10. The Case for No CPU Limits
  11. Static CPU Manager: Pinning Integer-CPU Pods
  12. Memory Manager: NUMA-Local Allocation
  13. Topology Manager: Hint Merging and Scopes
  14. ResourceQuota: Namespace Caps
  15. LimitRange: Per-Object Defaults and Bounds
  16. Quota Scopes and ScopeSelector
  17. PriorityClass and Preemption Recap
  18. Extended and Scalar Resources
  19. Hugepages: 2Mi and 1Gi
  20. Ephemeral Storage
  21. PID Limits and pid.available
  22. The Kubelet's Reservation Model: Allocatable
  23. Cgroup Hierarchy on a Kubernetes Node
  24. Eviction Signals and Thresholds
  25. Eviction Ranking: BestEffort First
  26. Eviction vs OOM Kill: Proactive vs Reactive
  27. Throttling Timeline and cpu.stat
  28. In-Place Pod Resize
  29. VPA Integration (Forward Ref)
  30. The Right-Sizing Workflow
  31. Observability: Metrics, Alerts, Dashboards
  32. Pitfalls
  33. TL;DR

1. Why Requests and Limits Exist

Linux gives processes everything they ask for, until it can't, at which point it kills somebody. That works for a laptop. It is catastrophic in a multi-tenant cluster where 100 pods on the same node compete for 64 cores and 256 GiB of RAM. The cluster needs two distinct controls, applied at two distinct times:

  1. A capacity reservation, consulted at scheduling time, that says "this pod needs at least N CPU and M memory; don't place it where that's already promised to someone else." This is requests.
  2. A hard ceiling, enforced at runtime by the kernel, that says "no matter how much spare capacity exists, this pod may not consume more than X CPU or Y memory." This is limits.

These are different problems with different consequences for setting them wrong:

  • A pod with no requests is treated by the scheduler as free — the node looks unloaded even when its actual cgroup usage is at 95%. The scheduler keeps packing pods on, the kernel starts evicting, and the on-call learns about the cascade at 03:00.
  • A pod with no limits can consume the entire node if nothing else is competing. Usually that is fine — the kernel will multiplex CPU fairly via cpu.weight (§4), and memory pressure will trigger eviction before the kernel OOM-kills the wrong thing (§24). But a runaway memory allocator with no memory.max will take the node down. Memory limits are mandatory; CPU limits often hurt more than they help (§10).
  • A pod with requests == limits (Guaranteed class, §3) is the simplest case: scheduler and kernel agree on the budget, no overcommit, no throttling, no eviction (until the node itself runs out of headroom).

The cluster economy depends on this asymmetry. Requests sum to what's promised; the cluster autoscaler grows the fleet when promised > total capacity. Limits sum to what could possibly be consumed; that number is usually >> total capacity (overcommit) and is what gives Kubernetes its bin-packing efficiency. The QoS class (§3) is just a label that summarizes the relationship between these two numbers for a pod, and the kubelet uses that label to decide who dies first.

1.1 The mental model in one diagram

       PROMISED                                          ACTUAL USE
       (requests)                                        (cgroup stats)
       ─────────                                         ──────────────
         │                                                    │
         │  scheduler reserves capacity                       │  kernel enforces limits
         │  on a Node based on this                           │  via cpu.max / memory.max
         │                                                    │  on this
         ▼                                                    ▼
   ┌──────────────────────────────────────────────────────────────────┐
   │  NODE: Allocatable = Capacity                                    │
   │                       - kubeReserved                             │
   │                       - systemReserved                           │
   │                       - hard-evictionThresholds                  │
   │                                                                  │
   │   ┌────────────────────────────────────────────────────────┐     │
   │   │  Σ(pod.requests)  ≤  Allocatable     (scheduler invariant)│  │
   │   │  Σ(pod.limits)    can EXCEED Allocatable (overcommit)    │   │
   │   └────────────────────────────────────────────────────────┘     │
   │                                                                  │
   │   When Σ(actual_usage) approaches Allocatable:                   │
   │     - cpu: cpu.weight arbitrates (no kill)                       │
   │     - mem: eviction manager picks a victim (proactive)           │
   │     - mem: kernel OOM picks a victim (reactive, last resort)     │
   └──────────────────────────────────────────────────────────────────┘

The rest of this chapter is the long form of that diagram.


2. The Two-Knob Model: Scheduling vs Enforcement

Every container.resources field in a Pod spec maps to exactly one of these two semantic categories. Memorize the table:

Field Read by When Effect
requests.cpu kube-scheduler scheduling counts against node's Allocatable.cpu
requests.cpu kubelet cgroup setup written as cpu.weight (proportional share)
requests.memory kube-scheduler scheduling counts against node's Allocatable.memory
requests.memory kubelet (MemQoS) cgroup setup written as memory.min (soft reservation)
limits.cpu scheduler (ignored) does NOT affect placement
limits.cpu kubelet cgroup setup written as cpu.max (quota/period throttling)
limits.memory scheduler (ignored) does NOT affect placement
limits.memory kubelet cgroup setup written as memory.max (kernel OOM on exceed)
requests.hugepages-2Mi scheduler scheduling matched to node's hugepages-2Mi capacity
requests.ephemeral-storage scheduler scheduling counts against node's ephemeral-storage cap
limits.ephemeral-storage kubelet eviction runtime triggers per-pod ephemeral eviction
requests.<extended> scheduler scheduling matches via Device Plugin (CDI/Allocate)

Two rules from that table that surprise people:

  • limits.cpu and limits.memory are invisible to the scheduler. A node already 100% scheduled by requests will not fit another pod even if limits would technically allow it; conversely, a node with 50% requests will accept another pod even if existing pods' limits add up to 400%. The scheduler is a promises engine, not a usage engine.
  • requests.cpu is not just a number — it becomes cpu.weight in the cgroup. Two pods on the same node with requests.cpu: 100m each will get equal CPU when contending. A pod with requests.cpu: 1000m gets 10x as much CPU as a pod with requests.cpu: 100m under contention. This is the only way requests influence runtime; there's no per-pod CPU floor.

The flow:

Pod manifest                           Cluster                          Node
────────────                           ───────                          ────

spec.containers[].resources.requests   ┌─────────────┐                  
   cpu: 250m                           │  scheduler  │                  
   memory: 512Mi   ───────────────►    │  Filter:    │                  
                                       │  NodeRes-   │                  
                                       │  ourcesFit  │                  
                                       │  plugin     │                  
                                       └──────┬──────┘                  
                                              │ Bind: spec.nodeName     
                                              ▼                         
                                                                  ┌──────────────┐
                                                                  │  kubelet     │
                                                                  │  pod worker  │
                                                                  └──────┬───────┘
                                                                         │
                                                                         ▼
                                                                  ┌──────────────┐
                                                                  │  cm.New      │
                                                                  │  PodContainer│
                                                                  │  Manager     │
                                                                  └──────┬───────┘
                                                                         │
                                                                         ▼
                                                          /sys/fs/cgroup/kubepods.slice/
                                                            kubepods-burstable.slice/
                                                              kubepods-burstable-pod<UID>.slice/
                                                                cri-containerd-<CID>.scope/
                                                                  cpu.weight     ← from requests.cpu
                                                                  cpu.max        ← from limits.cpu
                                                                  memory.min     ← from requests.memory (MemQoS)
                                                                  memory.high    ← from limits.memory * 0.8 (MemQoS)
                                                                  memory.max     ← from limits.memory
                                                                  pids.max       ← from podPidsLimit

The scheduler reads the top half of that picture. The kernel enforces the bottom half. The two halves never speak; their only contract is "the scheduler must not over-promise requests."

2.1 What if you omit fields?

Container spec Resulting QoS class What happens
nothing BestEffort first to die under pressure
requests.cpu only Burstable scheduled by CPU; no memory promise
requests.memory only Burstable scheduled by memory; CPU is share-only
requests.cpu, requests.memory (< limits) Burstable guaranteed only up to requests
requests == limits (both) Guaranteed scheduler & kernel align; no throttle
limits only (no requests) Burstable kubelet defaults requests = limits
LimitRange configured for namespace depends kubelet/apiserver fills defaults

The fourth-from-bottom row catches everybody: setting only limits does not make a pod Guaranteed — instead the apiserver's defaulter copies limits into requests, which does make it Guaranteed if no other container in the pod has different settings. We'll formalize this in §3.


3. QoS Class Derivation

The QoS class is computed once, by the apiserver/kubelet, and stored as pod.status.qosClass. It's not a user input. The decision tree is exactly this (pkg/apis/core/v1/helper/qos/qos.go):

                        ┌─────────────────────────────────────┐
                        │  For every container in the pod:    │
                        │  - requests AND limits set for      │
                        │    both CPU and memory?             │
                        │  - requests.cpu == limits.cpu?      │
                        │  - requests.mem == limits.mem?      │
                        └──────┬───────────────────────────┬──┘
                               │ yes for ALL                │ no for any
                               ▼                            ▼
                        ┌────────────────┐         ┌────────────────────────┐
                        │  Guaranteed    │         │  Any container has     │
                        │                │         │  any requests OR       │
                        │  oom_score_adj │         │  any limits set?       │
                        │  = -997        │         └─────┬──────────┬───────┘
                        └────────────────┘               │ yes      │ no
                                                         ▼          ▼
                                            ┌─────────────────┐  ┌──────────────┐
                                            │   Burstable     │  │  BestEffort  │
                                            │                 │  │              │
                                            │  oom_score_adj  │  │  oom_score   │
                                            │  computed per-  │  │  _adj = 1000 │
                                            │  container (§8) │  │              │
                                            └─────────────────┘  └──────────────┘

The exact rules in code:

// Simplified from pkg/apis/core/v1/helper/qos/qos.go (GetPodQOS).
func GetPodQOS(pod *v1.Pod) v1.PodQOSClass {
    requests := v1.ResourceList{}
    limits   := v1.ResourceList{}
    isGuaranteed := true
    zeroQuantity := resource.MustParse("0")

    for _, c := range pod.Spec.Containers {
        // Aggregate requests and limits across containers (sum).
        for name, q := range c.Resources.Requests {
            if name == v1.ResourceCPU || name == v1.ResourceMemory {
                existing := requests[name]
                existing.Add(q)
                requests[name] = existing
            }
        }
        for name, q := range c.Resources.Limits {
            if name == v1.ResourceCPU || name == v1.ResourceMemory {
                existing := limits[name]
                existing.Add(q)
                limits[name] = existing
            }
        }
        // Guaranteed requires: BOTH CPU and memory limits set,
        // AND limits == requests, on EVERY container.
        if len(c.Resources.Limits) == 0 ||
           c.Resources.Limits.Cpu().IsZero() ||
           c.Resources.Limits.Memory().IsZero() ||
           c.Resources.Requests.Cpu().Cmp(*c.Resources.Limits.Cpu()) != 0 ||
           c.Resources.Requests.Memory().Cmp(*c.Resources.Limits.Memory()) != 0 {
            isGuaranteed = false
        }
    }
    if isGuaranteed && len(pod.Spec.Containers) > 0 {
        return v1.PodQOSGuaranteed
    }
    if requests.Cpu().Cmp(zeroQuantity) == 0 &&
       requests.Memory().Cmp(zeroQuantity) == 0 &&
       limits.Cpu().Cmp(zeroQuantity) == 0 &&
       limits.Memory().Cmp(zeroQuantity) == 0 {
        return v1.PodQOSBestEffort
    }
    return v1.PodQOSBurstable
}

Three implications:

  • Init containers count. A Guaranteed pod must have requests == limits on every init container and every regular container. A single init container that omits limits.memory demotes the whole pod to Burstable.
  • Native sidecars (1.28+) count. Same rule applies. Adding an Istio sidecar with no resources downgrades a Guaranteed app to Burstable.
  • Ephemeral storage, hugepages, GPUs, and extended resources do not affect QoS. Only cpu and memory participate in the QoS computation.

3.1 Practical implications of each class

Class Eviction order OOM oom_score_adj CFS treatment NUMA pinning Use case
BestEffort first 1000 cpu.weight = 1 (min) never dev jobs, batch with no SLO
Burstable middle 2..999 (formula §8) cpu.weight from req never most stateless services
Guaranteed last -997 cpu.weight = MAX and fixed cpu.max if limit set yes (with static CPU manager + int CPUs) latency-sensitive, RT, low-latency DB

The single most important operational implication: never put production workloads in BestEffort. They are the first thing the eviction manager kills under any pressure, with no warning. A misconfigured Deployment with no resource block on a busy node will go into CrashLoopBackOff and you will spend two hours blaming the application before you check kubectl get pod -o jsonpath='{.status.qosClass}'.


3A. CPU Fundamentals: Logical vs Physical, CPU-Seconds, CFS, and Concurrency

Before diving into how Kubernetes maps requests.cpu and limits.cpu to cgroup knobs (§4), you need a precise understanding of what "1 CPU" actually means, how the kernel accounts for CPU time, and how the CFS scheduler distributes that time across competing processes and threads. Without this, every throttling bug, every VPA recommendation, and every capacity-planning spreadsheet is built on fuzzy intuition.

3A.1 What is a "CPU" in Kubernetes?

Kubernetes defines 1 CPU = 1 logical CPU as seen by the Linux kernel. This is the unit that appears in requests.cpu and limits.cpu. But "logical CPU" maps to different physical realities depending on the hardware:

┌─────────────────────────────────────────────────────────────────────────┐
│  PHYSICAL MACHINE                                                       │
│                                                                         │
│  Socket 0 (physical package)          Socket 1 (physical package)       │
│  ┌──────────────────────────┐        ┌──────────────────────────┐       │
│  │  Physical Core 0         │        │  Physical Core 8         │       │
│  │  ├─ Logical CPU 0 (HT0)  │        │  ├─ Logical CPU 16 (HT0) │       │
│  │  └─ Logical CPU 1 (HT1)  │        │  └─ Logical CPU 17 (HT1) │       │
│  │                          │        │                          │       │
│  │  Physical Core 1         │        │  Physical Core 9         │       │
│  │  ├─ Logical CPU 2 (HT0)  │        │  ├─ Logical CPU 18 (HT0) │       │
│  │  └─ Logical CPU 3 (HT1)  │        │  └─ Logical CPU 19 (HT1) │       │
│  │                          │        │                          │       │
│  │  ...                     │        │  ...                     │       │
│  │  Physical Core 7         │        │  Physical Core 15        │       │
│  │  ├─ Logical CPU 14 (HT0) │        │  ├─ Logical CPU 30 (HT0) │       │
│  │  └─ Logical CPU 15 (HT1) │        │  └─ Logical CPU 31 (HT1) │       │
│  └──────────────────────────┘        └──────────────────────────┘       │
│                                                                         │
│  Total: 2 sockets × 8 physical cores × 2 hyperthreads = 32 logical CPUs│
│  Kubernetes sees: Allocatable.cpu = 32 (minus kubeReserved)             │
└─────────────────────────────────────────────────────────────────────────┘

The terminology hierarchy:

Term Definition Example
Socket (package) A physical CPU chip in a motherboard slot. A 2-socket server has 2 chips.
Physical core An independent execution unit with its own ALU, FPU, L1/L2 cache. An Intel Xeon 8380 has 40 physical cores per socket.
Logical CPU (hardware thread) A schedulable execution context. With Hyper-Threading (Intel) or SMT (AMD), each physical core exposes 2 logical CPUs. Without HT/SMT, logical CPU = physical core. 40 cores × 2 HT = 80 logical CPUs per socket.
vCPU (cloud) What the cloud provider calls a logical CPU. AWS: 1 vCPU = 1 hyperthread. GCP: 1 vCPU = 1 hyperthread. Azure: 1 vCPU = 1 hyperthread (usually). An m5.4xlarge has 16 vCPUs = 16 hyperthreads = 8 physical cores.

Critical implication: Kubernetes cpu: 1 does not mean "one physical core." It means one hyperthread's worth of compute. Two hyperthreads on the same physical core share execution resources (ALU pipelines, L1 cache, TLB). A workload pinned to logical CPUs 0 and 1 (same physical core) gets roughly 1.0–1.3× the throughput of a single logical CPU, not 2×. The topology manager (§13) and static CPU manager (§11) exist precisely to give you control over this.

How to see what a node has:

# On the node itself:
$ lscpu
Architecture:          x86_64
CPU(s):                32          # ← this is what Kubernetes sees
On-line CPU(s) list:   0-31
Thread(s) per core:    2           # ← Hyper-Threading enabled
Core(s) per socket:    8
Socket(s):             2
NUMA node(s):          2
NUMA node0 CPU(s):     0-7,16-23
NUMA node1 CPU(s):     8-15,24-31

# From Kubernetes:
$ kubectl get node worker-1 -o jsonpath='{.status.capacity.cpu}'
32
$ kubectl get node worker-1 -o jsonpath='{.status.allocatable.cpu}'
31500m   # 32 - 500m kubeReserved

3A.2 CPU-Seconds: The Accounting Unit

The kernel doesn't think in "CPU cores" — it thinks in CPU-seconds (or microseconds internally). One CPU-second means one logical CPU was busy executing instructions for one second. This is the fundamental currency.

Definition: If a process runs on one logical CPU for T seconds of wall-clock time, and spends B seconds of that actively executing (not sleeping/waiting), it has consumed B CPU-seconds.

The math:

CPU-seconds consumed = Σ (time each logical CPU spent executing this cgroup's tasks)

CPU utilization (cores) = CPU-seconds consumed / wall-clock seconds elapsed

So when kubectl top pod shows CPU: 500m, it means: "over the last measurement window, this pod consumed 0.5 CPU-seconds per wall-clock second" — equivalent to keeping one logical CPU 50% busy, or two logical CPUs 25% busy each.

3A.2.1 Simple Mental Model: How CPU-Seconds Work

Think of CPU-seconds like man-hours on a construction site. If a job requires "160 man-hours", you can complete it with different combinations of workers (cores) and time:

  • Massive Parallelism (160 Cores): If 160 cores work at 100% capacity for just 1 real-world second, that equals 160 CPU-seconds (\(160 \text{ cores} \times 1 \text{ second} = 160\)).
  • Single Worker (1 Core): If you use only 1 core running at 100% capacity, it will take 160 real-world seconds to complete the same 160 CPU-seconds of work (\(1 \text{ core} \times 160 \text{ seconds} = 160\)).
  • Team of Workers (16 Cores): If you use 16 cores running at full speed at the same time, your task will finish in 10 real-world seconds, but it still consumes 160 CPU-seconds (\(16 \text{ cores} \times 10 \text{ seconds} = 160\)).
  • Partial Speed (4 Cores at 50%): If 4 cores run at 50% load for 20 real-world seconds, you consume 40 CPU-seconds (\(4 \text{ cores} \times 0.50 \text{ load} \times 20 \text{ seconds} = 40\)).
┌─────────────────────────────────────────────────────────────────────────────┐
│ REAL-WORLD TIME VS. CPU-SECONDS                                             │
│                                                                             │
│ Option 1: 1 Core @ 100% load for 160 seconds   ➜  1 × 160  = 160 CPU-sec   │
│ Option 2: 16 Cores @ 100% load for 10 seconds  ➜  16 × 10  = 160 CPU-sec   │
│ Option 3: 160 Cores @ 100% load for 1 second   ➜  160 × 1  = 160 CPU-sec   │
│ Option 4: 4 Cores @ 50% load for 20 seconds    ➜  4 × 0.5 × 20 = 40 CPU-sec│
└─────────────────────────────────────────────────────────────────────────────┘

Key Takeaway: Real-world time (wall-clock time) and CPU time are distinct. CPU-seconds = (Number of Cores Used) × (Wall-Clock Seconds) × (Average % Core Load).

3A.3 Worked Example: Single Pod, Single Container

Scenario:
  Pod A: requests.cpu=500m, limits.cpu=1
  Node: 4 logical CPUs
  Measurement window: 10 seconds
  Pod A runs a single-threaded web server.
  During the 10s window, the server processes requests that consume
  4.2 CPU-seconds of total CPU time.

Accounting:
  CPU-seconds consumed:  4.2
  Wall-clock elapsed:    10s
  Average CPU usage:     4.2 / 10 = 0.42 cores = 420m

  CFS quota check (per 100ms period):
    quota  = 1 core × 100ms = 100ms per period
    actual = 420m average → ~42ms per period on average
    Result: well under quota. No throttling.

  kubectl top pod:
    NAME     CPU(cores)
    pod-a    420m

  What the scheduler reserved:
    500m of node's Allocatable → 3500m remaining for other pods.

3A.4 Worked Example: Multiple Pods Competing Over 30 Seconds

Scenario:
  Node: 4 logical CPUs (4000m total compute per second)

  Pod A: requests.cpu=1000m, limits.cpu=2000m   (weight: 39,  quota: 200ms/100ms)
  Pod B: requests.cpu=500m,  limits.cpu=1000m   (weight: 20,  quota: 100ms/100ms)
  Pod C: requests.cpu=500m,  no limits          (weight: 20,  quota: unlimited)

  All three pods are CPU-bound (each trying to use as much CPU as possible)
  over a 30-second window.

Step 1: Total requested = 1000m + 500m + 500m = 2000m ≤ 4000m → scheduler places all three.

Step 2: Runtime — all pods are CPU-bound, so contention exists.
  Total weights: 39 + 20 + 20 = 79

  Without limits, CFS would distribute by weight:
    Pod A share: 39/79 × 4 cores = 1.97 cores
    Pod B share: 20/79 × 4 cores = 1.01 cores
    Pod C share: 20/79 × 4 cores = 1.01 cores

  But limits cap Pod A and Pod B:
    Pod A: min(1.97, 2.0) = 1.97 cores  ← under limit, no throttling
    Pod B: min(1.01, 1.0) = 1.00 core   ← right at limit, will start throttling
    Pod C: min(1.01, ∞)   = 1.01 cores  ← no limit, no throttling

  Leftover from Pod B's cap: 0.01 cores redistributed by weight to A and C.

  Final approximate allocation:
    Pod A: ~1.98 cores → 1.98 × 30 = 59.4 CPU-seconds in 30s
    Pod B: ~1.00 core  → 1.00 × 30 = 30.0 CPU-seconds in 30s
    Pod C: ~1.02 cores → 1.02 × 30 = 30.6 CPU-seconds in 30s
    Total:  4.00 cores → 120.0 CPU-seconds = 4 cores × 30s ✓

Step 3: What if Pod A goes idle at T=15?
  T=0 to T=15:  three-way contention (as above)
  T=15 to T=30: only Pod B and Pod C competing, 4 cores available
    Pod B: min(4 × 20/40, 1.0) = min(2.0, 1.0) = 1.0 core  ← still capped!
    Pod C: min(4 × 20/40, ∞)   = 2.0 cores                  ← takes all the slack
    Remaining 1.0 core: goes to Pod C (only uncapped consumer)
    Pod C actually gets: 3.0 cores (the entire remainder)

  This is why no-limits pods are better neighbors: they absorb slack.

3A.5 Multiprocessing vs Multithreading: How CFS Sees Them

From the kernel's perspective, the CFS scheduler schedules tasks (kernel-level threads). Whether your application uses multiprocessing (separate PIDs, separate address spaces) or multithreading (same PID, shared address space) doesn't matter for CPU accounting — both produce kernel tasks that CFS schedules independently.

The critical difference is how they interact with cgroups:

┌────────────────────────────────────────────────────────────────────────────┐
│ MULTIPROCESSING (e.g., Python multiprocessing, Gunicorn prefork)          │
│                                                                            │
│  Container cgroup: cri-containerd-<CID>.scope                             │
│  │                                                                        │
│  ├── PID 1 (entrypoint / master)     ← task 1                            │
│  ├── PID 2 (worker process 1)        ← task 2, own address space          │
│  ├── PID 3 (worker process 2)        ← task 3, own address space          │
│  └── PID 4 (worker process 3)        ← task 4, own address space          │
│                                                                            │
│  All 4 PIDs are in the SAME cgroup.                                       │
│  cpu.max applies to the SUM of all 4 PIDs' CPU time.                      │
│  Each process has its own memory (RSS); total counts against memory.max.  │
│                                                                            │
│  CPU accounting: 4 workers × 100ms wall-clock = up to 400ms CPU-time      │
│  If cpu.max = "200000 100000" (limit 2 cores):                            │
│    4 workers can collectively use 200ms per 100ms period.                 │
│    Each worker gets ~50ms on average → effectively 0.5 cores each.        │
└────────────────────────────────────────────────────────────────────────────┘

┌────────────────────────────────────────────────────────────────────────────┐
│ MULTITHREADING (e.g., Java, Go goroutines, C++ std::thread)               │
│                                                                            │
│  Container cgroup: cri-containerd-<CID>.scope                             │
│  │                                                                        │
│  └── PID 1 (JVM / Go runtime)        ← main task                         │
│      ├── TID 1 (main thread)         ← task 1                            │
│      ├── TID 2 (worker thread 1)     ← task 2, shared address space       │
│      ├── TID 3 (worker thread 2)     ← task 3, shared address space       │
│      ├── TID 4 (GC thread)           ← task 4, shared address space       │
│      └── TID 5 (I/O thread)          ← task 5, shared address space       │
│                                                                            │
│  All TIDs are in the SAME cgroup (threads inherit parent's cgroup).       │
│  cpu.max applies to the SUM of all TIDs' CPU time — identical to above.   │
│  Memory is shared; RSS is counted once (not per-thread).                  │
│                                                                            │
│  CPU accounting: same as multiprocessing — CFS doesn't care.              │
│  5 threads × 100ms wall-clock = up to 500ms CPU-time                      │
│  If cpu.max = "200000 100000" (limit 2 cores):                            │
│    5 threads collectively use 200ms per 100ms period.                     │
│    Each thread gets ~40ms on average → effectively 0.4 cores each.        │
└────────────────────────────────────────────────────────────────────────────┘

Key insight: CFS quota is a per-cgroup aggregate across all tasks (threads and processes). Whether you fork 4 processes or spawn 4 threads, the cgroup consumes CPU-seconds at the same rate. The quota doesn't care about your concurrency model.

3A.6 The Multi-Threaded Quota Trap: A Concrete Example

This is the single most common production surprise with CPU limits. Let's walk through it step by step:

Scenario:
  Java application with 8 request-handling threads.
  Each request takes 5ms of CPU time.
  Requests arrive uniformly: 100 requests/second.
  Pod spec: limits.cpu = 500m → cpu.max = "50000 100000" (50ms per 100ms period)

Per-period accounting:
  In a 100ms period, ~10 requests arrive.
  If each request is handled by a separate thread simultaneously:
    8 threads × 5ms each = 40ms CPU-time consumed in ~5ms wall-clock
    Quota used: 40ms out of 50ms → OK, no throttling.

  But what if a brief burst of 15 requests arrives in one period?
    8 threads handle the first 8 simultaneously:
      8 × 5ms = 40ms CPU-time in ~5ms wall-clock. Quota remaining: 10ms.
    Next 7 requests start, but after ~1.25ms wall-clock (7 × 1.25 ≈ 8.75ms CPU):
      Quota exhausted! All 8 threads are FROZEN for the rest of the period.

  Timeline:
  ├─ 0ms ──── 5ms ──── 6.25ms ──── THROTTLED ──── 100ms ─┤
    [8 req]     [7 req]   ↑                                
                          quota exhausted                   
                          93.75ms of enforced idle!         

  Those 7 requests each took 5ms of CPU but ~95ms of wall-clock.
  From the application's perspective: p99 latency jumped from 5ms to 95ms.
  From the node's perspective: 7 other cores were completely idle.

  This is the CFS quota trap. The solution: remove limits.cpu (§10)
  or increase it to cover burst parallelism (limits.cpu ≥ 8 × 5ms/100ms = 400m
  per thread × 8 threads = 3200m for zero throttling under full parallelism).

3A.7 CFS Internals: How the Scheduler Actually Works

The Completely Fair Scheduler (CFS) is the default Linux task scheduler (since kernel 2.6.23). Understanding its internals explains why CPU accounting, weights, and quotas behave the way they do.

How CFS distributes CPU time (the weight mechanism)

CFS maintains a per-CPU red-black tree of runnable tasks, sorted by virtual runtime (vruntime). The key idea:

vruntime_increment = actual_runtime / weight

A task with higher weight accumulates vruntime slower → stays toward the
left of the tree → gets picked to run more often.

When the scheduler needs to pick a task, it picks the leftmost node (lowest vruntime). After running for a time slice, the task's vruntime is updated and it's re-inserted. The effect is that tasks with higher weight naturally get more CPU time proportional to their weight.

In the Kubernetes context, cpu.weight (derived from requests.cpu) feeds directly into this mechanism:

Pod A: requests.cpu=1000m → cpu.weight=39
Pod B: requests.cpu=250m  → cpu.weight=10

Under contention on a single CPU:
  vruntime grows 39/10 = 3.9× slower for Pod A's tasks
  → Pod A gets ~3.9× more CPU time than Pod B
  → Pod A: 39/(39+10) = 79.6% of the CPU
  → Pod B: 10/(39+10) = 20.4% of the CPU

With no contention:
  Both can run whenever they want. vruntime is irrelevant.
  Pod B can use 100% of an idle CPU.

How CFS enforces quotas (the bandwidth controller)

The CFS bandwidth controller (CONFIG_CFS_BANDWIDTH) is the mechanism behind cpu.max. It operates at the cgroup level, not per-task:

Data structures (simplified from kernel/sched/fair.c):

struct cfs_bandwidth {
    u64 quota;          // max CPU-time per period (from cpu.max)
    u64 period;         // period length in ns (default 100ms)
    u64 runtime;        // remaining runtime in current period
    int nr_throttled;   // count of throttled runqueues
    s64 throttled_time; // total throttled nanoseconds
};

The lifecycle per period:

┌──────────────────────────────────────────────────────────────────────┐
│  Period starts: runtime = quota (e.g., 50ms for limits.cpu=500m)    │
│                                                                      │
│  CFS picks tasks from this cgroup to run on various CPUs:           │
│                                                                      │
│  CPU 0: task runs 12ms → runtime = 50 - 12 = 38ms                  │
│  CPU 2: task runs  8ms → runtime = 38 -  8 = 30ms                  │
│  CPU 0: task runs 15ms → runtime = 30 - 15 = 15ms                  │
│  CPU 1: task runs 10ms → runtime = 15 - 10 =  5ms                  │
│  CPU 3: task runs  5ms → runtime =  5 -  5 =  0ms                  │
│                                                                      │
│  runtime == 0: THROTTLE                                              │
│  All tasks in this cgroup are dequeued from all CPUs' runqueues.    │
│  They cannot be scheduled until the next period.                     │
│                                                                      │
│  ... time passes ... no tasks from this cgroup run ...               │
│                                                                      │
│  Period ends: runtime is reset to quota. Tasks are re-enqueued.     │
│  nr_throttled++. throttled_time += (time spent throttled).           │
└──────────────────────────────────────────────────────────────────────┘

The subtle point: runtime is global across all CPUs. If your cgroup has tasks running on 4 CPUs simultaneously, it burns through quota at 4× the wall-clock rate. This is exactly why the multi-threaded trap (§3A.6) happens.

CFS scheduling in a two-level cgroup hierarchy

Kubernetes pods live in a nested cgroup hierarchy (§6). CFS handles this via hierarchical scheduling:

kubepods.slice/ (weight: 39)
├── kubepods-burstable.slice/ (weight: 33)
│   ├── pod-A/ (weight: 39, i.e., requests.cpu=1)
│   │   └── container-A/ (weight: 39)
│   │       ├── thread-1
│   │       └── thread-2
│   └── pod-B/ (weight: 20, i.e., requests.cpu=500m)
│       └── container-B/ (weight: 20)
│           └── thread-1
└── kubepods-besteffort.slice/ (weight: 1)
    └── pod-C/ (weight: 1)
        └── container-C/ (weight: 1)
            └── thread-1

Under full contention on a 4-core node:

  Level 1: kubepods.slice gets all 4 cores (only pod cgroup at this level).

  Level 2: burstable vs besteffort
    burstable weight: 33
    besteffort weight: 1
    burstable gets: 33/34 × 4 = 3.88 cores
    besteffort gets: 1/34 × 4 = 0.12 cores

  Level 3 (within burstable): pod-A vs pod-B
    pod-A weight: 39
    pod-B weight: 20
    pod-A gets: 39/59 × 3.88 = 2.57 cores
    pod-B gets: 20/59 × 3.88 = 1.31 cores

  Result:
    Pod A (req 1000m): ~2570m  ← more than requested, absorbing slack
    Pod B (req 500m):  ~1310m  ← more than requested
    Pod C (req 0):     ~120m   ← scraps, but it asked for nothing

The hierarchy ensures that QoS classes are respected structurally, not just by weight values. BestEffort pods are children of a low-weight parent, so they get proportionally less even if their individual weight were higher (which it isn't — it's 1).

3A.8 GOMAXPROCS, JVM Thread Pools, and the Container CPU Problem

Many language runtimes auto-detect the number of available CPUs to size their thread pools. On bare metal, this works perfectly. Inside a container, it's a footgun:

Problem:
  Host has 64 logical CPUs.
  Container has limits.cpu=2 → cpu.max = "200000 100000".

  Go runtime: GOMAXPROCS defaults to runtime.NumCPU() = 64
    → 64 goroutines can run in parallel
    → burns through 200ms quota in ~3ms wall-clock
    → throttled for 97ms

  JVM: Runtime.getRuntime().availableProcessors() = 64
    → ForkJoinPool.commonPool size = 63
    → GC parallel threads = ~16
    → same quota burn problem

  Python: os.cpu_count() = 64
    → multiprocessing.Pool() defaults to 64 workers
    → same problem, but per-process

  Node.js: os.cpus().length = 64
    → cluster.fork() in a loop = 64 workers
    → same problem

Solution:
  These runtimes should read the cgroup limit, not /proc/cpuinfo.

  Go 1.19+:    Automatically reads cpu.max. GOMAXPROCS = ceil(quota/period).
               With limits.cpu=2 → GOMAXPROCS=2. Correct.
               For older Go or when you want explicit control:
               Use uber-go/automaxprocs: import _ "go.uber.org/automaxprocs"

  JVM 10+:     Reads cpu.max via -XX:+UseContainerSupport (default on).
               availableProcessors() = ceil(quota/period).
               JVM 8u191+ also supports this.
               Override: -XX:ActiveProcessorCount=2

  Python:      os.cpu_count() still returns 64 (as of 3.13).
               Use: len(os.sched_getaffinity(0)) for cpuset-aware count,
               or manually read /sys/fs/cgroup/cpu.max.

  Node.js:     os.cpus().length still returns 64.
               Manually set cluster workers: Math.min(os.cpus().length, N).
               Or read the cgroup limit.

3A.9 CPU-Seconds in Prometheus: Connecting the Theory to Observability

The cAdvisor metric container_cpu_usage_seconds_total is a monotonically increasing counter of CPU-seconds consumed. Everything in this section connects to it:

# Instantaneous CPU usage in cores (= CPU-seconds per second):
rate(container_cpu_usage_seconds_total{container="myapp"}[5m])

# This number means:
#   0.5  → the container uses 500m (half a logical CPU) on average
#   2.0  → the container uses 2000m (two logical CPUs) on average
#   0.05 → the container uses 50m on average

# Compare to requests to see if right-sized:
rate(container_cpu_usage_seconds_total{container="myapp"}[5m])
 /
kube_pod_container_resource_requests{resource="cpu", container="myapp"}

# > 1.0 means the pod is using more than it requested (bursting)
# < 0.3 means the pod is over-provisioned (wasting scheduler capacity)
# 0.5–0.8 is the healthy range for latency-sensitive services

3A.10 Summary Table: CPU Concepts at a Glance

Concept What it is Kubernetes relevance
Physical core Independent execution unit on the die Not directly visible to K8s; matters for cache/NUMA (§13)
Logical CPU Schedulable hardware thread (with HT: 2 per physical core) This is what cpu: 1 means
CPU-second 1 logical CPU busy for 1 second The unit of container_cpu_usage_seconds_total
cpu.weight CFS proportional share (1–10000) Derived from requests.cpu; arbitrates under contention
cpu.max CFS bandwidth quota (quota/period μs) Derived from limits.cpu; hard cap regardless of idle CPUs
CFS period Time window for quota accounting (default 100ms) Shorter = less throttle latency, more overhead
Throttling Cgroup frozen until next period after quota exhausted The reason multi-threaded apps get surprise latency spikes
GOMAXPROCS / thread pool size Runtime's concurrency level Must match cgroup limit, not host CPU count

3B. Capacity Metrics: Usage, Utilization, Core-Hours, and Memory Accounting

When you move from "make my pod run" to "how much infrastructure does my team/service/cluster actually consume, and are we right-sized?", you need precise definitions of usage, utilization, and cost units — for both CPU and memory. These terms are used loosely in conversation, precisely in billing systems, and inconsistently across monitoring tools. Getting them wrong leads to chargebacks that are unfair, autoscalers that thrash, and capacity plans that are fiction.

3B.1 CPU: Usage vs Utilization vs Allocation — Three Different Numbers

These three terms are often confused but measure fundamentally different things:

┌────────────────────────────────────────────────────────────────────────────┐
│ CPU ALLOCATION (requests)                                                  │
│   What the scheduler reserved. A bookkeeping number.                      │
│   Does NOT change with actual load. Fixed at pod creation.                │
│   Unit: cores (milliCPU)                                                  │
│   Source: kube_pod_container_resource_requests{resource="cpu"}             │
│                                                                            │
│ CPU USAGE                                                                  │
│   How many CPU-seconds the cgroup actually consumed per wall-clock second.│
│   This IS the actual load. Changes constantly.                            │
│   Unit: cores (CPU-seconds per second)                                    │
│   Source: rate(container_cpu_usage_seconds_total[5m])                      │
│                                                                            │
│ CPU UTILIZATION                                                            │
│   Usage expressed as a fraction of some reference.                         │
│   Which reference? That's where the confusion lives.                       │
│   Unit: percentage (0–100% or 0–N00% for multi-core)                      │
│   Source: depends on which denominator you pick (see below)               │
└────────────────────────────────────────────────────────────────────────────┘

The denominator problem — "utilization of what?":

Metric Formula What it answers Range
Utilization vs request usage / requests.cpu "Is this pod right-sized?" 0–∞ (can exceed 100% if bursting)
Utilization vs limit usage / limits.cpu "How close to throttling?" 0–100% (can't exceed if CFS enforced)
Utilization vs node usage / node.allocatable.cpu "What fraction of this node is this pod consuming?" 0–100%
Utilization vs cluster Σ(usage) / Σ(node.allocatable.cpu) "How loaded is the entire cluster?" 0–100%

HPA uses utilization-vs-request by default (type: Resource, target.averageUtilization). This is critical: if your pod requests 1000m but uses 200m, HPA sees 20% utilization and may scale down — even though the pod might be doing useful work. If your pod requests 100m but uses 500m (bursting with no limits), HPA sees 500% and scales up aggressively.

Prometheus queries for each

# 1. Raw CPU usage in cores
rate(container_cpu_usage_seconds_total{
    container!="", container!="POD",
    namespace="prod"
}[5m])

# 2. Utilization vs request (what HPA sees)
rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
  /
kube_pod_container_resource_requests{resource="cpu"}

# 3. Utilization vs limit
rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
  /
kube_pod_container_resource_limits{resource="cpu"}

# 4. Node-level utilization (all pods on the node)
sum by (node) (rate(container_cpu_usage_seconds_total{container!=""}[5m]))
  /
kube_node_status_allocatable{resource="cpu"}

# 5. Cluster-wide utilization
sum(rate(container_cpu_usage_seconds_total{container!=""}[5m]))
  /
sum(kube_node_status_allocatable{resource="cpu"})

# 6. Allocation efficiency: how much of what's reserved is actually used
sum(rate(container_cpu_usage_seconds_total{container!=""}[5m]))
  /
sum(kube_pod_container_resource_requests{resource="cpu"})

3B.2 CPU Core-Hours: The Billing Unit

Cloud providers and internal chargeback systems bill CPU by the core-hour: one logical CPU running for one hour.

Definition:

1 core-hour = 1 logical CPU × 1 hour = 3600 CPU-seconds

There are two very different ways to calculate core-hours, and which one you pick determines who pays what:

Method 1: Allocation-based (what you reserved)

core-hours_allocated = Σ (requests.cpu × pod_uptime_hours)

Example:
  Pod A: requests.cpu=2, ran for 24 hours
  Pod B: requests.cpu=500m, ran for 8 hours
  Pod C: requests.cpu=100m, ran for 24 hours

  Total = (2 × 24) + (0.5 × 8) + (0.1 × 24)
        = 48 + 4 + 2.4
        = 54.4 core-hours allocated

PromQL:

# Core-hours allocated over the last 24h, per namespace
sum by (namespace) (
    kube_pod_container_resource_requests{resource="cpu"}
) * 24  # hours

# More precise: integrate over actual pod uptime using recording rules
# (since pods come and go throughout the day)
sum by (namespace) (
    increase(kube_pod_container_resource_requests{resource="cpu"}[24h:5m]) * (5/60)
)

Method 2: Usage-based (what you actually consumed)

core-hours_used = Σ (CPU-seconds consumed) / 3600

Example (same pods, measured usage):
  Pod A: used avg 1.2 cores over 24 hours → 1.2 × 24 = 28.8 core-hours
  Pod B: used avg 450m over 8 hours       → 0.45 × 8 = 3.6 core-hours
  Pod C: used avg 30m over 24 hours       → 0.03 × 24 = 0.72 core-hours

  Total = 28.8 + 3.6 + 0.72
        = 33.12 core-hours used

PromQL:

# Core-hours consumed over the last 24h, per namespace
sum by (namespace) (
    increase(container_cpu_usage_seconds_total{container!="", container!="POD"}[24h])
) / 3600

Why the gap matters

Allocation-based: 54.4 core-hours
Usage-based:      33.12 core-hours
Efficiency:       33.12 / 54.4 = 60.9%
Waste:            39.1% of reserved CPU sat idle

The 39.1% gap is real infrastructure cost. Every idle reserved core blocks the scheduler from placing other pods. This gap is what VPA (§29) and right-sizing (§30) aim to close.

Billing model Pros Cons When to use
Allocation-based Simple, predictable bills; teams know their cost upfront; encourages right-sizing (you pay for waste) Punishes bursty workloads; teams game it by under-requesting Internal chargeback, capacity planning
Usage-based Fair; you pay for what you consume; no waste penalty Unpredictable bills; doesn't account for reserved capacity that others can't use; encourages over-requesting Cloud billing (AWS/GCP on-demand), dev environments
Blended Max(allocation, usage) or weighted average; captures both reservation cost and actual consumption More complex to explain and implement Large orgs with mature FinOps

3B.3 Memory: Usage, Working Set, and RSS — Not the Same Thing

Memory accounting is more complex than CPU because the kernel tracks multiple overlapping counters, and picking the wrong one leads to wrong eviction decisions, wrong OOM thresholds, and wrong capacity plans.

┌────────────────────────────────────────────────────────────────────────────┐
│  MEMORY COUNTERS IN A CGROUP (cgroup v2)                                   │
│                                                                            │
│  memory.current                                                            │
│    Total memory charged to this cgroup. Includes:                         │
│    ├── RSS (Resident Set Size): anonymous pages (heap, stack, mmap)        │
│    ├── Page cache: file-backed pages (read/written files, mmap'd files)   │
│    ├── Kernel memory: slab, page tables, socket buffers                   │
│    ├── tmpfs / shmem: shared memory segments                              │
│    └── Swap (if enabled): pages swapped out but still charged             │
│                                                                            │
│  memory.stat (selected fields):                                            │
│    anon          = RSS (anonymous pages only)                              │
│    file          = page cache (file-backed pages only)                     │
│    kernel        = kernel memory (slab + page tables + ...)                │
│    shmem         = tmpfs / POSIX shared memory                             │
│    inactive_anon = RSS pages not recently accessed                         │
│    inactive_file = page cache pages not recently accessed                  │
│                                                                            │
│  DERIVED METRICS:                                                          │
│    Working Set Size (WSS) = memory.current - inactive_file                 │
│      This is what Kubernetes reports as "memory usage"                     │
│      (via cAdvisor → metrics-server → kubectl top)                        │
│                                                                            │
│    RSS = memory.stat.anon                                                  │
│      Just the anonymous (non-file-backed) pages                           │
│                                                                            │
│    Cache = memory.stat.file                                                │
│      File-backed pages that the kernel CAN drop under pressure            │
└────────────────────────────────────────────────────────────────────────────┘

Why the distinction matters enormously:

Metric What it includes Can the kernel reclaim it? What happens at memory.max?
memory.current Everything: RSS + cache + kernel + shmem Partially (cache is reclaimable) Kernel reclaims cache first, then OOM kills
Working Set (current - inactive_file) RSS + active cache + kernel + shmem Mostly not (these are "needed" pages) This is what should be compared to limits
RSS (memory.stat.anon) Only heap/stack/anonymous mmap No (must kill the process) RSS alone exceeding memory.max → guaranteed OOM
Cache (memory.stat.file) Only file-backed pages Yes (kernel drops them freely) Cache exceeding limit → kernel just drops pages, no OOM

The critical insight: A pod can have memory.current = 3 GiB and memory.max = 2 GiB simultaneously and not be OOM-killed — if 1.5 GiB is reclaimable page cache. The kernel reclaims cache pages as needed. But kubectl top reports working set, which excludes inactive file pages, so you see ~1.5–2 GiB — a more honest number.

Corner case: I/O-heavy workloads

Scenario:
  Pod runs a log processor that reads 100 GiB of files per hour.
  requests.memory=512Mi, limits.memory=1Gi.

  What happens:
    The kernel caches recently-read file pages in the cgroup's page cache.
    memory.current quickly reaches 1Gi (memory.max).
    BUT: most of those pages are file cache (reclaimable).
    memory.stat.anon (RSS) is only 200Mi (the actual application heap).

    The kernel continuously evicts old cache pages to make room for new ones.
    No OOM. No eviction. The pod is perfectly healthy.

    BUT: `container_memory_usage_bytes` (memory.current) shows 1Gi.
    An uninformed alert on memory.current > 90% of limit fires constantly.

  Solution: alert on working_set, not memory.current:
    container_memory_working_set_bytes < limits.memory * 0.9

3B.4 Memory Byte-Hours: The Memory Billing Unit

The memory equivalent of a core-hour is a GiB-hour (or MiB-hour): one GiB of memory held for one hour.

Definition:

1 GiB-hour = 1 GiB of memory reserved (or used) for 1 hour
           = 1,073,741,824 bytes × 3600 seconds
           = 3,865,470,566,400 byte-seconds

Allocation-based (reserved)

GiB-hours_allocated = Σ (requests.memory_GiB × pod_uptime_hours)

Example:
  Pod A: requests.memory=4Gi, ran 24 hours  → 4 × 24 = 96 GiB-hours
  Pod B: requests.memory=512Mi, ran 8 hours → 0.5 × 8 = 4 GiB-hours
  Pod C: requests.memory=256Mi, ran 24 hours → 0.25 × 24 = 6 GiB-hours

  Total = 106 GiB-hours allocated

PromQL:

# GiB-hours allocated over 24h, per namespace
sum by (namespace) (
    kube_pod_container_resource_requests{resource="memory"}
) / (1024^3) * 24

Usage-based (consumed)

For memory, use working set as the usage metric (not memory.current, which inflates with reclaimable cache):

# GiB-hours consumed over 24h, per namespace (using working set)
sum by (namespace) (
    avg_over_time(container_memory_working_set_bytes{
        container!="", container!="POD"
    }[24h])
) / (1024^3) * 24
Example (same pods, measured working set):
  Pod A: avg working set 2.8Gi over 24h → 2.8 × 24 = 67.2 GiB-hours
  Pod B: avg working set 400Mi over 8h  → 0.39 × 8 = 3.12 GiB-hours
  Pod C: avg working set 100Mi over 24h → 0.098 × 24 = 2.35 GiB-hours

  Total = 72.67 GiB-hours used
  Efficiency = 72.67 / 106 = 68.6%
  Waste = 31.4% of reserved memory sat unused

3B.5 Why Memory Efficiency Is Harder to Optimize Than CPU

CPU is compressible: if you over-request, the slack is used by other pods via cpu.weight. The cores aren't wasted — other pods absorb them.

Memory is incompressible: if you reserve 4 GiB but use 2 GiB, those 2 GiB of Allocatable are blocked on the scheduler's books. No other pod can be placed using that headroom (the scheduler checks Σ(requests) ≤ Allocatable, not Σ(actual)). The memory sits physically available but logically locked.

4-GiB node, 3 pods:

  Scheduler's view (requests):      Kernel's view (actual RSS):
  ┌────────────────────────┐        ┌────────────────────────┐
  │ Pod A: req 2Gi         │        │ Pod A: using 800Mi     │
  │ Pod B: req 1Gi         │        │ Pod B: using 600Mi     │
  │ Pod C: req 1Gi         │        │ Pod C: using 400Mi     │
  ├────────────────────────┤        ├────────────────────────┤
  │ Total: 4Gi (FULL)      │        │ Total: 1.8Gi (45%)     │
  │ No more pods accepted! │        │ 2.2Gi physically free! │
  └────────────────────────┘        └────────────────────────┘

  The scheduler says the node is full. The kernel disagrees.
  This 55% gap is pure waste — but reducing requests risks OOM.

This is why memory right-sizing is both more important and more dangerous than CPU right-sizing: - Reduce requests too much → scheduler packs more pods → a traffic spike pushes actual usage above what the node has → OOM kills cascade. - Reduce limits too much → any memory spike kills the pod. Memory doesn't throttle — it kills. - CPU over-request → just wastes scheduler capacity. No crashes. Other pods still burst into the slack.

3B.6 Corner Cases and Gotchas

Corner case 1: Init containers inflate core-hours

Pod spec:
  initContainers:
    - name: db-migrate
      resources:
        requests:
          cpu: "4"          # needs 4 cores for a 30-second migration
          memory: "8Gi"
  containers:
    - name: web
      resources:
        requests:
          cpu: "500m"       # steady-state needs
          memory: "1Gi"

The scheduler computes effective request as:
  cpu: max(4, 0.5) = 4     (init and regular don't run simultaneously)
  memory: max(8, 1) = 8Gi

But for allocation-based billing:
  If the pod runs for 24 hours, you're billed for 4 cores × 24h = 96 core-hours.
  The init container ran for 30 seconds. The web container uses 500m.
  Actual usage: (4 × 30/3600) + (0.5 × 24) ≈ 0.033 + 12 = 12.03 core-hours.
  Efficiency: 12.5%!

Solution: use restartPolicy: Never on init containers (1.28+ sidecar containers),
or split the migration into a separate Job with its own resource requests.

Corner case 2: JVM heap vs container memory

Pod: limits.memory=2Gi
JVM: -Xmx1536m (1.5Gi max heap)

Expected: JVM uses ≤1.5Gi, container is safe.
Reality: JVM uses ~1.5Gi heap + ~400Mi off-heap (metaspace, thread stacks,
         direct buffers, native code, JIT compiler) = ~1.9Gi.
         Plus kernel memory (page tables, socket buffers) = ~2.0Gi.

Result: OOM at random times. "But I set Xmx to 1.5Gi!"

Rule of thumb:
  limits.memory ≥ Xmx + (0.5 × Xmx)  for JVM workloads
  Or use -XX:MaxRAMPercentage=75 (JVM 10+): JVM reads cgroup limit
  and sets heap to 75% of it, leaving 25% for off-heap.

Corner case 3: Page cache doesn't count toward requests but does toward limits

Pod: requests.memory=1Gi, limits.memory=4Gi
App: reads a 3Gi dataset via mmap.

memory.current = 3.5Gi (1Gi RSS + 2.5Gi page cache)
memory.stat.anon = 1Gi
container_memory_working_set_bytes = 2Gi (active pages)

Scheduler reserved 1Gi. Node has room for 3 more 1Gi pods.
But the kernel charged 3.5Gi to this cgroup's memory.current.
If memory.current approaches memory.max (4Gi), kernel reclaims
page cache pages — application slows down (I/O instead of cache hit)
but doesn't die.

But if the application also grows RSS to 2Gi (leak):
  memory.current = 4.5Gi → exceeds memory.max = 4Gi
  Kernel tries to reclaim 0.5Gi of cache pages.
  If only 0.5Gi cache remains and it can't reclaim enough → OOM.

Bottom line: page cache is a hidden memory consumer that's
usually harmless but can push you into OOM during memory leaks.

Corner case 4: Measuring short-lived pods

Job pods that run for 30 seconds:
  - rate() over 5m averages the usage over 5 minutes
  - A pod that used 4 cores for 30 seconds shows as:
    rate(container_cpu_usage_seconds_total[5m]) = 4 × 30 / 300 = 0.4 cores
  - This is mathematically correct but operationally misleading.

For billing: use increase() instead of rate():
  increase(container_cpu_usage_seconds_total[1h]) / 3600 = core-hours
  This correctly captures the 30-second burst.

For capacity: track the PEAK concurrent usage, not the average:
  max_over_time(
    sum(rate(container_cpu_usage_seconds_total{namespace="batch"}[1m]))[1h:1m]
  )

Corner case 5: CPU steal and noisy neighbors on shared cloud instances

Your pod runs on an EC2 m5.xlarge (4 vCPUs, shared tenancy).
Kubernetes sees 4 allocatable cores.
Your pod requests 2 cores.

But the hypervisor is oversubscribing the physical host.
Your "4 vCPUs" are actually shares on a 128-core machine
running 40 other VMs.

Symptom:
  container_cpu_usage_seconds_total shows 1.8 cores.
  But actual throughput is 40% lower than on a dedicated host.
  rate(node_cpu_seconds_total{mode="steal"}[5m]) > 0.05 (5% stolen).

What happened:
  The hypervisor took CPU cycles from your VM to serve other tenants.
  Kubernetes has NO visibility into this — cAdvisor doesn't see steal time
  at the container level (only at the node level via node_exporter).

Your CFS quota accounting is correct (you used 1.8 cores of cgroup time),
but the WALL-CLOCK throughput is degraded because each "core-second"
delivered less actual work.

Solution: monitor node_cpu_seconds_total{mode="steal"} and alert if > 5%.
Consider dedicated/metal instances for latency-sensitive workloads.

3B.7 The Capacity Planning Spreadsheet: Putting It All Together

Here's how to calculate whether your cluster has enough capacity and when to scale:

Cluster: 10 nodes × 32 logical CPUs = 320 cores total
kubeReserved: 500m per node → 5 cores reserved
Allocatable: 315 cores

Current state (from Prometheus):
  Σ(requests.cpu) across all pods:      220 cores (allocation)
  Σ(actual CPU usage) across all pods:  145 cores (usage)
  Σ(limits.cpu) across all pods:        480 cores (overcommit)

Derived metrics:
  Allocation ratio:     220 / 315 = 69.8%  (scheduler thinks 70% full)
  Utilization ratio:    145 / 315 = 46.0%  (actual node load is 46%)
  Allocation efficiency: 145 / 220 = 65.9% (pods use 66% of what they asked for)
  Overcommit ratio:     480 / 315 = 152%   (limits sum to 1.5× capacity)

Capacity signals:
  Allocation > 80% → add nodes (scheduler can't place new pods)
  Utilization > 70% → contention risk (CFS weight arbitration kicks in)
  Efficiency < 50%  → pods are over-requesting (VPA opportunity)
  Overcommit > 200% → high risk of cascading throttling under load

Memory (parallel calculation):
  Σ(requests.memory):  800 GiB allocated
  Σ(working_set):      520 GiB actual
  Allocatable memory:  960 GiB (10 × 96Gi)
  Allocation ratio:    83.3% — TIGHT. Scheduler will reject pods soon.
  Utilization ratio:   54.2% — plenty of physical headroom.
  Efficiency:          65.0% — 35% waste in reservation.

  Action: either right-size memory requests (risky) or add nodes (safe).

3B.8 Summary Table: All Capacity Metrics at a Glance

Metric CPU version Memory version Formula Use case
Allocation Σ requests.cpu (cores) Σ requests.memory (GiB) From pod specs Scheduler capacity, chargeback
Usage rate(cpu_usage_seconds_total) (cores) working_set_bytes (GiB) From cgroup stats Actual load, autoscaling
Utilization (vs request) usage / requests.cpu working_set / requests.memory Usage ÷ Allocation HPA, right-sizing
Utilization (vs capacity) usage / allocatable.cpu working_set / allocatable.memory Usage ÷ Node cap Node saturation
Core-hours / GiB-hours Σ(requests.cpu × hours) Σ(requests.memory × hours) Allocation × time Billing, cost reports
Efficiency usage / allocation working_set / allocation How much of reservation is used FinOps, waste detection
Overcommit Σ limits / allocatable Σ limits / allocatable Limits ÷ Capacity Risk assessment

4. CPU Semantics: Fractional Requests, cpu.weight, cpu.max

CPU in Kubernetes is fractional and normalized to one logical core. One core's worth of CPU is 1 or equivalently 1000m (milliCPU). Half a core is 500m. Two cores are 2 or 2000m. The unit "logical core" is whatever the kernel sees — on a 64-thread Xeon, that's 64. The scheduler doesn't know about hyperthreads vs physical cores until you reach the topology manager (§13).

4.1 How requests.cpu becomes cpu.weight

On cgroup v2, the relevant file is cpu.weight. Range: 1..10000, default 100. It is a proportional share: under contention, two tasks with weights w1 and w2 get CPU in ratio w1:w2; with no contention, both can run flat-out.

The kubelet's conversion (pkg/kubelet/cm/helpers_linux.go, MilliCPUToShares then translated to v2 weight by cri-containerd / the runtime):

v1 cpu.shares  = max(2, milliCPU * 1024 / 1000)    // legacy
v2 cpu.weight  = ((cpu.shares - 2) * 9999) / 262142 + 1
                 // then clamped to [1, 10000]

Concretely:

requests.cpu cpu.shares (v1) cpu.weight (v2)
100m 102 4
250m 256 10
500m 512 20
1 1024 39
2 2048 78
4 4096 157
8 8192 314
16 16384 626

What this means in practice: if pod A has requests.cpu: 1 (weight 39) and pod B has requests.cpu: 100m (weight 4) and they're on the same fully-loaded core, A gets ~90% of the core, B gets ~10%. With no contention (e.g., A is idle), B can use the entire core.

4.2 How limits.cpu becomes cpu.max

On cgroup v2, cpu.max has the form "$QUOTA $PERIOD" in microseconds:

$ cat /sys/fs/cgroup/kubepods.slice/.../cri-containerd-abc.scope/cpu.max
50000 100000

That reads as: "in each 100ms period, this cgroup may use 50ms of CPU time, summed across all CPUs." The default period is 100ms (configurable via --cpu-cfs-quota-period). The kubelet's formula:

quota  = limits.cpu_milli * period / 1000
       = 500 * 100000 / 1000
       = 50000   (microseconds per period, for limits.cpu=500m)

The two interesting cases:

  • limits.cpu: 100m → cpu.max = "10000 100000". The container can use 10ms of CPU per 100ms period. If it tries to use more (e.g., spawns 4 threads each running flat-out), it gets throttled — see §9.
  • limits.cpu: 2 → cpu.max = "200000 100000". The container can use 200ms of CPU per 100ms period, i.e., two cores' worth, achievable only by running two threads in parallel.
  • No limits.cpu → cpu.max = "max 100000". No throttling. The container can saturate the whole machine if cpu.weight lets it.

4.3 Why cpu.weight and cpu.max are separate knobs

The two interact in a non-obvious way:

Scenario: pod A has requests.cpu=1, limits.cpu=2.
         pod B has requests.cpu=1, no limits.
         Node has 4 cores.

         Both pods are CPU-bound, running 4 threads each.

                                  cpu.weight   cpu.max
         pod A:                       39        200000 / 100000
         pod B:                       39        max

                                  Result
         ────────────────────────────────────────────────────────
         pod A capped at 2 cores by CFS quota (cpu.max).
         pod B uses the remaining 2 cores (weights equal, A maxed).
         No throttling on B. A's throttle counter increments steadily.

So cpu.weight arbitrates between cgroups; cpu.max is a per-cgroup absolute cap. Knowing this is the difference between debugging tail latency in 5 minutes and 5 hours.

4.4 The 1m floor

You can request as little as 1m (one milliCPU). The kubelet rejects anything finer. The corresponding cpu.weight is 1 (after clamping). Sub-milli precision is meaningless because CFS scheduling granularity is in microseconds and the period is 100ms — you literally cannot account for less than one part in 100000.


5. Memory Semantics: Bytes, Reservation, memory.max

Memory in Kubernetes is bytes, with the usual unit suffixes (Ki, Mi, Gi, Ti, plus power-of-ten K, M, G, T). One important gotcha: M (1 000 000 bytes) is not Mi (1 048 576 bytes). Most production specs use Mi/Gi.

5.1 requests.memory is reserved at scheduling

Unlike requests.cpu, requests.memory doesn't have an obvious cgroup mapping. The scheduler reserves bytes against node.status.allocatable.memory, but the kernel does not receive a per-pod floor by default. If three pods on a node each requested 1 GiB and one of them is using 4 GiB while the others use 100 MiB, the kernel is happy until the node's total free memory drops past the kubelet's eviction threshold — at which point the kubelet picks a victim (§24).

With the MemoryQoS feature gate (§7) enabled, the kubelet does write memory.min = requests.memory to give the kernel a hint not to reclaim from that cgroup. But MemoryQoS is still beta in 1.33 and disabled by default in most distributions.

5.2 limits.memory is a hard ceiling: memory.max

This one is enforced by the kernel, immediately, and brutally. On cgroup v2:

$ cat /sys/fs/cgroup/kubepods.slice/.../memory.max
536870912

Means: this cgroup may not use more than 512 MiB of memory. The instant the cgroup's memory.current would exceed memory.max, the kernel either:

  1. Reclaims memory inside the cgroup (page-cache pages, swap if allowed).
  2. If reclaim fails, invokes the cgroup OOM killer, which picks the worst-scoring process inside this cgroup and kills it.

The kill is a SIGKILL to the chosen victim. The kubelet observes the death via the container runtime's exit code (137 = 128 + 9), and marks the container's lastState.terminated.reason = "OOMKilled". If the container's restartPolicy permits, it's restarted.

5.3 Why memory limits are mandatory in production

CPU you can leave open (§10). Memory you cannot. A single misbehaving allocator with no memory.max will:

  1. Eat all RSS on the node.
  2. Push the node into eviction territory (§24).
  3. The kubelet evicts pods to recover.
  4. Eviction may not be fast enough: a malloc loop can burn 10 GiB/s on modern hardware.
  5. The kernel global OOM killer fires.
  6. The kernel global OOM killer does not know about QoS or the kubelet's preferences directly; it consults oom_score_adj (§8), which the kubelet did set per pod, but the kernel may still pick the wrong process if everything has the same score.
  7. Worst case: the kernel kills the kubelet, the container runtime, or systemd. Node goes NotReady.

Setting memory.max puts a fence around each pod: a runaway allocator dies inside its own cgroup before it threatens the node.

5.4 Swap

By default, Kubernetes (since 1.8) requires swap to be disabled (swapoff -a). Otherwise the kubelet refuses to start. The reason: swap interacts badly with memory.max accounting and with HPA's memory-based scaling. Since 1.28 there's an experimental failSwapOn=false and a NodeSwap feature gate that allows --swap-behavior=LimitedSwap, but it's beta and not widely deployed. Treat swap as off.


6. The cgroup-v2 Tree Under kubepods.slice

The kubelet builds a four-level cgroup hierarchy on every node. Knowing the shape of this tree is what makes debugging tractable — most "weird resource bugs" can be answered with cat on the right file.

/sys/fs/cgroup/                                             ← cgroup-v2 unified root
├── init.scope/                                             ← PID 1 (systemd or sysvinit)
├── system.slice/                                           ← systemd-managed services
│   ├── kubelet.service/                                    ← the kubelet itself
│   ├── containerd.service/                                 ← the CRI runtime
│   └── ...
├── user.slice/                                             ← interactive logins
│
└── kubepods.slice/                                         ← LEVEL 1: ALL pods
    │   cpu.weight = 39        # weight 1000 of node CPU
    │   memory.max = max       # uncapped; node-level only
    │   cpu.max = max
    │
    ├── kubepods-besteffort.slice/                          ← LEVEL 2: BestEffort QoS
    │   │   cpu.weight = 1     # lowest priority share
    │   │   memory.max = max
    │   │
    │   └── kubepods-besteffort-pod<UID1>.slice/            ← LEVEL 3: one pod
    │       │   cpu.weight = 1
    │       │   memory.max = max
    │       │
    │       ├── cri-containerd-<CID-pause>.scope/           ← LEVEL 4: pause container
    │       │      pids.current = 1
    │       │
    │       └── cri-containerd-<CID-app>.scope/             ← LEVEL 4: app container
    │              cpu.weight = 1
    │              cpu.max = max
    │              memory.max = max
    │              pids.max = 4096      # podPidsLimit, default
    │
    ├── kubepods-burstable.slice/                           ← LEVEL 2: Burstable QoS
    │   │   cpu.weight = 33    # share between node CPU
    │   │   memory.max = max
    │   │
    │   └── kubepods-burstable-pod<UID2>.slice/             ← LEVEL 3: one pod
    │       │   cpu.weight = sum(container.requests.cpu)
    │       │   memory.max = sum(container.limits.memory) or max
    │       │
    │       ├── cri-containerd-<CID-pause>.scope/
    │       ├── cri-containerd-<CID-init>.scope/            ← init container (terminated)
    │       └── cri-containerd-<CID-app>.scope/
    │              cpu.weight = container.requests.cpu derived
    │              cpu.max    = container.limits.cpu derived (or "max")
    │              memory.max = container.limits.memory (or "max")
    │              memory.min = container.requests.memory (MemQoS only)
    │              memory.high = limits.memory * (throttlingFactor)
    │
    └── kubepods-pod<UID3>.slice/                           ← LEVEL 3 (no L2 for Guaranteed!)
        │   # Guaranteed pods are direct children of kubepods.slice
        │   cpu.weight = sum(container.requests.cpu)
        │   memory.max = sum(container.limits.memory)
        │
        └── cri-containerd-<CID-app>.scope/
               cpuset.cpus = 4-7    # set by static CPU manager (§11)
               cpuset.mems = 0      # set by memory manager (§12)

A few non-obvious things:

  • There is no kubepods-guaranteed.slice. Guaranteed pods hang directly off kubepods.slice/kubepods-pod<UID>.slice. This is intentional: putting them under a shared parent would impose a parent cpu.weight that splits CPU between the parent slice and its siblings. By being direct children, Guaranteed pods get their full proportional share against the node root.
  • The pause container exists in its own scope. It's the namespace anchor (chapter 11). Its cgroup is mostly empty (1 PID, ~100 KiB).
  • The CRI runtime names scopes cri-containerd-<containerID>.scope or crio-<containerID>.scope depending on the runtime. containerd uses the former, CRI-O the latter. The kubelet itself doesn't write these; the CRI runtime does, with knobs derived from the OCI runtime spec (config.json → linux.resources → unified map for cgroup-v2).

6.1 Reading the tree in practice

$ systemctl status kubelet | grep CGroup
   CGroup: /kubepods.slice

$ ls /sys/fs/cgroup/kubepods.slice/ | head
cgroup.controllers
cgroup.events
cgroup.freeze
cgroup.max.depth
cgroup.max.descendants
cgroup.procs
cgroup.subtree_control
cgroup.threads
cgroup.type
cpu.idle
cpu.max
cpu.pressure
cpu.stat
cpu.weight
kubepods-besteffort.slice
kubepods-burstable.slice
kubepods-podb1234...slice               ← Guaranteed pod
kubepods-podc5678...slice               ← Guaranteed pod
memory.current
memory.events
memory.high
memory.max
memory.min
memory.pressure
memory.stat
pids.current
pids.max

$ cat /sys/fs/cgroup/kubepods.slice/cpu.weight
33

$ cat /sys/fs/cgroup/kubepods.slice/memory.current
4831838208                  # ~4.5 GiB currently used by ALL pods

This tree is the only truth about what's actually happening on the node. Prometheus's cAdvisor walks it every 30 seconds; kubectl top reads from metrics-server which reads from /metrics/resource on the kubelet which reads from cAdvisor.


7. memory.high and the MemoryQoS Feature Gate

memory.max is a hard limit: hit it, get killed. memory.high (introduced in cgroup-v2) is a soft limit: hit it, get throttled — the kernel forces direct reclaim inside the offending cgroup, slowing it down so it has less time to allocate. No kill, just back-pressure.

The MemoryQoS feature gate (beta in 1.27, still default-off in 1.33) makes the kubelet write memory.high automatically:

memory.min  = container.requests.memory
memory.high = floor(throttlingFactor * (limit - request) + request)
memory.max  = container.limits.memory

Where throttlingFactor defaults to 0.9. So a container with requests.memory: 100Mi, limits.memory: 1Gi:

memory.min  = 100 MiB     # kernel won't reclaim below this if it can help it
memory.high = 100 + 0.9 * (1024 - 100) = 100 + 832 = 932 MiB
memory.max  = 1024 MiB    # OOM here

The behavior:

  • Below 100 MiB: kernel preserves these pages (won't push to swap if disabled, won't drop page-cache aggressively).
  • 100–932 MiB: normal accounting.
  • 932 MiB–1024 MiB: kernel forces reclaim within this cgroup whenever the cgroup tries to allocate. Allocations slow down. memory.events:high counter increments.
  • 1024 MiB: cgroup OOM killer fires.

Why this matters: without MemoryQoS, an aggressive memory grower just hits memory.max and dies. With MemoryQoS, the kernel pushes back before the cliff, giving the process a chance to slow down, complete a transaction, run a finalizer, or shed load. Apps that respect MEMORY_PRESSURE (via memory.pressure PSI) can adapt.

The cost: memory.high throttling is implemented by stalling the allocating task. If your app is single-threaded and the allocator is the hot path, throttling looks like a latency spike. So MemoryQoS is great for tail-latency steadiness (no OOMs) but trades against tail-latency minimum (no throttle stalls).

The feature gate:

# kubelet flag:
--feature-gates=MemoryQoS=true
# In KubeletConfiguration:
featureGates:
  MemoryQoS: true

7.1 Inspecting the PSI signals

cgroup-v2 exposes memory.pressure, cpu.pressure, io.pressure per cgroup. These are PSI (Pressure Stall Information) counters:

$ cat /sys/fs/cgroup/kubepods.slice/kubepods-burstable.slice/.../memory.pressure
some avg10=0.42 avg60=0.18 avg300=0.05 total=438219
full avg10=0.00 avg60=0.00 avg300=0.00 total=0

some = the percentage of time at least one task in this cgroup was stalled waiting for memory. full = the percentage of time all tasks were stalled. avg10 etc. are decaying averages over 10/60/300 seconds. total is monotonic microseconds of stall.

In a Prometheus alert:

- alert: PodMemoryPressureHigh
  expr: rate(container_memory_pressure_seconds_total{level="some"}[5m]) > 0.10
  for: 10m
  labels:
    severity: warning
  annotations:
    summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} memory PSI > 10%"
    description: "Container is stalling on memory allocation > 10% of the time. Consider raising memory request/limit or enabling MemoryQoS."

8. OOM Scoring: Per-QoS oom_score_adj

When the kernel needs to kill something to free memory, it scores every process and kills the highest scorer. The score is roughly:

score = (RSS_in_pages / total_pages) * 1000 + oom_score_adj
range: -1000 .. 1000

oom_score_adj lives in /proc/<pid>/oom_score_adj and is settable per-process. The kubelet writes it for every container process based on QoS class (pkg/kubelet/qos/policy.go):

// pkg/kubelet/qos/policy.go (simplified)
const (
    KubeletOOMScoreAdj         int = -999     // kubelet itself
    DockerOOMScoreAdj          int = -999     // (legacy)
    KubeProxyOOMScoreAdj       int = -999
    GuaranteedOOMScoreAdj      int = -997
    BestEffortOOMScoreAdj      int = 1000
)

func GetContainerOOMScoreAdjust(pod *v1.Pod, container *v1.Container,
                                memoryCapacity int64) int {
    if types.IsCriticalPod(pod) {
        // static pods, kube-system priorityClassName=system-cluster-critical
        return KubeletOOMScoreAdj
    }
    switch v1qos.GetPodQOS(pod) {
    case v1.PodQOSGuaranteed:
        return GuaranteedOOMScoreAdj
    case v1.PodQOSBestEffort:
        return BestEffortOOMScoreAdj
    }
    // Burstable: scale by how much memory the container is requesting
    // relative to node capacity.
    memReq := container.Resources.Requests.Memory().Value()
    oomScoreAdj := 1000 - (1000*memReq)/memoryCapacity
    if oomScoreAdj < BurstableOOMScoreAdjMin {
        oomScoreAdj = BurstableOOMScoreAdjMin   // 2
    }
    if oomScoreAdj >= BestEffortOOMScoreAdj {
        oomScoreAdj = BestEffortOOMScoreAdj - 1 // 999
    }
    return int(oomScoreAdj)
}

So a Burstable pod requesting 8 GiB on a 64 GiB node:

oom_score_adj = 1000 - (1000 * 8) / 64 = 1000 - 125 = 875

Higher score = killed first. A pod requesting more memory gets a lower score (less likely to be killed) — the reasoning is "this pod negotiated for more, so we should respect that more."

8.1 Score reference table

Process / class oom_score_adj Interpretation
kubelet -999 kernel will basically never kill it
container runtime (containerd, crio) -999 same
system-node-critical static pods -997 as critical as Guaranteed
Guaranteed pods -997 killed only if absolutely nothing else
Burstable, big memory request ~100-500 depends on request/capacity ratio
Burstable, tiny memory request 800-999 nearly as low priority as BestEffort
Burstable, no memory request, only CPU 999 scored as if no memory was requested
BestEffort pods 1000 first to die

The kubelet writes these scores by passing them to the CRI runtime in ContainerConfig.linux.resources.oomScoreAdj. The runtime sets them on the container's init process; the kernel inherits them to children (most of the time — see pitfalls).

8.2 What actually happens during OOM

Time   Event                                                                  
─────  ──────────────────────────────────────────────────────────────────────  
T+0    A container in pod X allocates 100 MiB more, exceeding its memory.max  
T+0    Kernel: memory.events:oom_kill incremented for cgroup                   
T+1ms  Kernel: scans cgroup.procs, computes oom_score for each PID            
T+2ms  Kernel: picks highest scorer, sends SIGKILL                            
T+3ms  Container PID 1 dies. All its children die (PID namespace teardown).   
T+5ms  Runtime (containerd) observes process death via exit pipe / waitpid    
T+10ms Runtime emits "Exit" event to CRI consumers (kubelet)                  
T+20ms PLEG observes container state change (or evented PLEG: immediate)      
T+50ms Kubelet's syncLoop runs SyncPod, sees container exited, exit code 137  
T+60ms Kubelet's statusManager PATCHes /pod/status:                           
        lastState.terminated.reason = "OOMKilled"                              
        lastState.terminated.exitCode = 137                                    
T+100ms If restartPolicy=Always, kubelet starts a new container (with backoff)

The 50–100ms gap between kernel kill and pod-status update is why dashboards sometimes show "container running" right after an OOM — you're looking at the pre-kill snapshot.


9. CFS Quota and CPU Throttling

CFS = Completely Fair Scheduler. It's the default Linux process scheduler. Its "quota" mechanism enforces cpu.max. The mechanism is exactly as described in §4.2: in every period (default 100ms), the cgroup gets at most quota microseconds of total CPU time across all CPUs.

9.1 The throttling timeline

limits.cpu = 400m  →  cpu.max = "40000 100000"  (40ms quota per 100ms period)

Single-threaded workload doing 30ms of work every 100ms:

Period N:    [work 30ms ───────][idle 70ms ──────────────────]   ← uses 30/40 quota
Period N+1:  [work 30ms ───────][idle 70ms ──────────────────]   ← uses 30/40 quota
No throttling. Throughput: 30ms work / 100ms = 30% of a core.    

Single-threaded workload doing 50ms of work every 100ms (busy):

Period N:    [work 40ms ───────────────────────][THROTTLED────]  ← hit quota at 40ms
                                                  exits early    
Period N+1:  [work 40ms ───────────────────────][THROTTLED────]  ← same
Throughput: 40ms work / 100ms = 40% of a core.                   
Tail latency: every burst > 40ms takes at least 100ms wall-clock.

Four-threaded workload, each thread doing 15ms of work every 100ms:

Period N:    [4 threads × 15ms parallel = 60ms quota in ~15ms wall]
             [THROTTLED for remaining 85ms ─────────────────────]
Throughput per thread: 15ms work / 100ms = 15%, but bursty.       
Latency for any single request: high (waiting for next period).  

That third scenario is the multi-threaded CFS throttling trap, and it's the single most common subtle latency bug in production Kubernetes. A Java app with 4 threads, each doing brief CPU work, can blow through a 40ms quota in 10ms of wall-clock and then wait 90ms for the next period — even though the node is 90% idle.

9.2 cpu.stat: the throttling counters

$ cat /sys/fs/cgroup/kubepods.slice/.../cri-containerd-abc.scope/cpu.stat
usage_usec      4231090112
user_usec       3892140000
system_usec     338950112
nr_periods      482103
nr_throttled    47281
throttled_usec  892341000
  • nr_periods: how many CFS periods the cgroup has been observed in. ~10 per second.
  • nr_throttled: how many of those periods ended with the cgroup hitting its quota and being throttled.
  • throttled_usec: total microseconds the cgroup spent in throttled state.

throttled_usec / (nr_periods * period_usec) = fraction of wall-clock time throttled.

The cAdvisor metric:

container_cpu_cfs_throttled_periods_total
container_cpu_cfs_throttled_seconds_total
container_cpu_cfs_periods_total

A useful alert:

- alert: PodCPUThrottlingHigh
  expr: |
    (rate(container_cpu_cfs_throttled_periods_total{container!="",container!="POD"}[5m])
     /
     rate(container_cpu_cfs_periods_total{container!="",container!="POD"}[5m])) > 0.25
  for: 15m
  labels:
    severity: warning
  annotations:
    summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} CPU throttled >25% of periods"
    description: |
      Container '{{ $labels.container }}' has been throttled in more than 25% of CFS
      periods for 15 minutes. Either raise limits.cpu, remove the limit entirely, or
      investigate why it's bursting (GC pause, async I/O completion storm, etc).

9.3 The kubelet's --cpu-cfs-quota-period

The CFS period defaults to 100ms. The kubelet flag --cpu-cfs-quota-period (since 1.18) lets you change it cluster-wide. Shorter periods (e.g., 10ms) reduce worst-case throttle wait time (you wait at most 10ms instead of 100ms) at the cost of higher accounting overhead. Long-running batch workloads benefit from longer periods; latency-sensitive apps benefit from shorter.

In KubeletConfiguration:

cpuCFSQuotaPeriod: "10ms"
cpuCFSQuota: true     # default; set false to disable CFS quota enforcement entirely

--cpu-cfs-quota=false is the nuclear option: it disables cpu.max enforcement for all containers on the node. The kubelet still writes cpu.weight so contention is resolved, but no quota throttling occurs. Useful for latency-critical clusters where you've decided requests-only is the right model (§10).

9.4 The fix that's not a fix: cpu.cfs_quota_us=-1 per container

Pre-1.18, an annotation on Borg-style clusters set cpu.cfs_quota_us=-1 (no quota) for specific pods. Kubernetes never adopted this per-pod opt-out. Either set --cpu-cfs-quota=false cluster-wide, or omit limits.cpu per pod, or raise the limit until throttling stops.


10. The Case for No CPU Limits

Stating it plainly: most workloads run faster without limits.cpu. The corollary is that limits.memory is still mandatory. This is the prevailing wisdom at companies that have measured it (Google, Buoyant, Lyft, Indeed, Bryan Boreham's now-famous KubeCon talk, etc.).

10.1 The argument

  1. CFS quota is enforced even when there's idle CPU on the node. A pod with limits.cpu: 500m will be throttled at 500m even if the other 31.5 cores are idle.
  2. cpu.weight from requests.cpu already prevents a runaway pod from starving its neighbors. If pod A has requests.cpu: 100m and pod B has requests.cpu: 1, under contention B gets 10x A's CPU. Without contention, both can use the whole machine.
  3. Multi-threaded apps (Java, Node, Go, anything with thread pools or async I/O completion) routinely have brief CPU bursts way above their average. CFS quota turns these bursts into tail-latency disasters.
  4. Memory is not multiplexable. If you give two pods a hard limit, the kernel can divide CPU fairly through time slicing. It cannot divide bytes fairly without picking a loser. Memory limits exist because memory is binary: you have it or you don't.

10.2 The counter-argument

  1. Predictable tenants are easier to schedule. A pod whose limits track its peak makes capacity planning trivial; without limits, you reason about worst-case-possible vs typical, which is uncomfortable.
  2. Hostile workloads — anything you can't trust to be well-behaved — should not have unlimited CPU. A crypto-mining container will saturate the node.
  3. Charge-back models often bill per limit. If you offer customers "you can use up to 4 vCPUs", they expect to be able to use 4 and you don't want them stealing 8.
  4. Resource pools that depend on bin-packing. HPA with target CPU utilization scales based on current_usage / requests. If pods routinely use more than requests (because no limit), the metric is dishonest; HPA may not scale when it should.

10.3 The compromise that works

In practice the production pattern that works across most stacks:

  • requests.cpu set to ~p50 of measured usage (or "what you need for stable steady-state"). This drives scheduling and cpu.weight.
  • No limits.cpu for latency-sensitive services. Let cpu.weight arbitrate.
  • limits.cpu only for batch/CI/dev workloads where you genuinely want to cap.
  • requests.memory and limits.memory both set, with limits.memory ≈ 1.5 × requests.memory. Limit protects the node; the margin between request and limit is your safety buffer.
  • Use ResourceQuota at the namespace level (§14) to cap aggregate CPU; this is where the "untrusted tenant" concern is handled correctly.

The cluster-wide opt-out (--cpu-cfs-quota=false) is appropriate when you've decided no workload should ever be CPU-throttled. The per-pod opt-out is just "leave limits.cpu blank."


11. Static CPU Manager: Pinning Integer-CPU Pods

For workloads where even occasional CFS throttling is unacceptable (low-latency trading, RT video transcoding, in-memory databases), the kubelet can pin specific containers to specific physical CPUs and forbid anyone else from running there. This is the static CPU manager (pkg/kubelet/cm/cpumanager/).

11.1 Activation

KubeletConfiguration:

cpuManagerPolicy: "static"
cpuManagerPolicyOptions:
  full-pcpus-only: "true"          # don't split hyperthread pairs
  align-by-socket: "true"          # don't span sockets
kubeReserved:
  cpu: "500m"
  memory: "1Gi"
systemReserved:
  cpu: "500m"
  memory: "1Gi"

Critical: switching from none to static requires draining the node and deleting /var/lib/kubelet/cpu_manager_state (the state file). The kubelet refuses to start if the state file is inconsistent.

11.2 What qualifies for pinning

A container is eligible for exclusive CPU pinning only if:

  1. The pod's QoS class is Guaranteed.
  2. The container's requests.cpu is an integer (1, 2, 4 — not 500m, not 1500m).

Containers that don't qualify (BestEffort, Burstable, fractional Guaranteed) run in the shared pool = (all CPUs) − (CPUs allocated to Guaranteed integer pods) − (reserved CPUs for the kubelet/system).

11.3 Allocation algorithm

The static manager keeps a CPU topology map (read from /sys/devices/system/cpu/): - Sockets → NUMA nodes → physical cores → logical CPUs (hyperthreads).

When a Guaranteed integer-CPU pod arrives:

1. Topology Manager (§13) asks the CPU manager: 
   "for N CPUs, give me a hint about which NUMA node alignments are feasible."
2. CPU manager computes: 
   - Prefer whole physical cores (both hyperthread siblings together).
   - Prefer alignment to a single NUMA node.
   - Prefer alignment to a single socket.
3. Topology Manager merges with hints from Memory Manager (§12) and Device 
   Manager, picks an aligned NUMA node.
4. CPU manager allocates the chosen CPUs, writes them to the container's 
   cpuset.cpus.
5. The kubelet ALSO updates the shared pool: removes those CPUs from every 
   non-pinned container's cpuset.cpus.

After:

# Pinned Guaranteed container
$ cat /sys/fs/cgroup/kubepods.slice/kubepods-pod<UID>.slice/.../cpuset.cpus
4-7

# Shared-pool container (Burstable, in another pod)
$ cat /sys/fs/cgroup/kubepods.slice/kubepods-burstable.slice/.../cpuset.cpus
0-3,8-31           ← 4-7 excluded

The kernel respects cpuset.cpus strictly: a process pinned to CPUs 4–7 will run only there. The kernel scheduler doesn't have to multiplex.

11.4 The state file

$ cat /var/lib/kubelet/cpu_manager_state
{
  "policyName": "static",
  "defaultCpuSet": "0-3,8-31",
  "entries": {
    "abc-pod-uid": {
      "container-1": "4-5",
      "container-2": "6-7"
    }
  },
  "checksum": 1234567
}

The checksum ensures the kubelet detects state corruption. If the checksum mismatches (e.g., manual edits, partial write), the kubelet refuses to start. Recovery: drain, delete file, restart, drain again, re-admit pods.

11.5 When static CPU manager hurts

  • Many small Burstable pods + a few Guaranteed integer pods = fragmented shared pool. The Burstable pods are squeezed onto fewer CPUs and contend more.
  • Mixed-workload nodes (e.g., 70% Burstable, 30% Guaranteed) often perform worse with static CPU manager than with none because the latency-sensitive Burstable pods can no longer span all cores during bursts.
  • Make sure the workload actually benefits before enabling. Benchmark.

12. Memory Manager: NUMA-Local Allocation

The memory manager (pkg/kubelet/cm/memorymanager/) does for memory what the CPU manager does for CPU: aligns allocations to NUMA nodes.

12.1 Activation

memoryManagerPolicy: "Static"     # or "None"
reservedMemory:
- numaNode: 0
  limits:
    memory: "1Gi"
- numaNode: 1
  limits:
    memory: "1Gi"

Only Static is currently supported (None is the default no-op).

12.2 What it does

For Guaranteed pods (only — same eligibility as CPU manager), the kubelet writes cpuset.mems to a NUMA mask:

$ cat /sys/fs/cgroup/kubepods.slice/kubepods-pod<UID>.slice/.../cpuset.mems
0

That tells the kernel: this container's memory allocations should come from NUMA node 0 only. Combined with the CPU manager pinning to CPUs on NUMA node 0, the result is local memory access — every load from this container hits the close DRAM controller, not the remote one (~2x faster).

12.3 The two failure modes

  • Topology fragmentation: pod requests 8 GiB and 4 CPUs, but no single NUMA node has both 8 GiB free and 4 free CPUs. With topology-manager policy single-numa-node, the kubelet rejects the pod (event: TopologyAffinityError). The scheduler will retry on another node — but if every node has the same fragmentation, the pod is permanently Pending.
  • Reservation mismatch: reservedMemory must sum to ≥ kubeReserved.memory + systemReserved.memory + evictionHard.memory.available. The kubelet refuses to start if this invariant is violated.

13. Topology Manager: Hint Merging and Scopes

The topology manager (pkg/kubelet/cm/topologymanager/) is the arbiter that asks the CPU manager, memory manager, and device manager for hints about NUMA placement, merges them into a single decision, and either admits or rejects the pod.

13.1 The hint-merge algorithm

                  ┌──────────────────────────────────────┐
                  │  Pod admitted to node                │
                  │  Pod has resources: cpu=4, mem=8Gi,  │
                  │                     nvidia.com/gpu=1 │
                  └─────────────────┬────────────────────┘
                                    │
              ┌─────────────────────┼──────────────────────┐
              │                     │                      │
              ▼                     ▼                      ▼
        ┌──────────┐         ┌───────────┐         ┌─────────────┐
        │ CPU Mgr  │         │ Memory    │         │ Device Mgr  │
        │ "I can   │         │ "I can    │         │ "GPU is on  │
        │ give you │         │ give you  │         │ NUMA 1; can │
        │ NUMA 0  │         │ NUMA 0    │         │ allocate    │
        │ or NUMA1 │         │ or NUMA1  │         │ only there" │
        │ aligned" │         │ aligned"  │         │             │
        └─────┬────┘         └─────┬─────┘         └──────┬──────┘
              │                    │                      │
              └──────────────┬─────┴──────────────────────┘
                             ▼
                  ┌──────────────────────┐
                  │  Topology Manager    │
                  │  merge:              │
                  │   NUMA 0 ∩ NUMA 1 =  │
                  │   { NUMA 1 } (GPU    │
                  │     forces it)       │
                  └──────────┬───────────┘
                             │
              ┌──────────────┼──────────────┐
              │  policy:     │              │
              │  - none:     │ accept anyway, no alignment
              │  - best-effort: prefer NUMA 1, accept anything
              │  - restricted: must use ONLY NUMA 1; reject if can't
              │  - single-numa-node: ONLY a single NUMA node; reject if cross
              └──────────────┼──────────────┘
                             ▼
                  ┌──────────────────────┐
                  │  decision: NUMA 1    │
                  │  Tell CPU Mgr: CPUs  │
                  │    on NUMA 1         │
                  │  Tell Memory Mgr:    │
                  │    cpuset.mems=1     │
                  │  Tell Device Mgr:    │
                  │    GPU on NUMA 1     │
                  └──────────────────────┘

13.2 The four policies

Policy Behavior
none (default) No hint merging. Each manager allocates independently.
best-effort Prefer aligned hint. If no aligned hint exists, accept anyway.
restricted Require aligned hint. Reject if every alignment crosses NUMA boundaries.
single-numa-node Require single-NUMA hint. Reject if any resource spans NUMA.

13.3 Scopes

Topology can be evaluated at two scopes:

  • container (default): per-container alignment. A pod with two containers can land each on a different NUMA node.
  • pod: all containers must align to the same NUMA node. Stricter; more likely to reject.
topologyManagerPolicy: "single-numa-node"
topologyManagerScope:  "pod"

13.4 Why this is hard

Topology decisions are made at pod admission time on the node, after the scheduler has already chosen the node. If the topology manager rejects the pod, it goes back to scheduler as Pending. The scheduler may pick the same node again (because it doesn't model NUMA), and the cycle repeats — TopologyAffinityError events accumulate, the pod never starts.

Mitigations: - Use the NUMA-aware scheduling plugin (in-tree, behind feature gate NodeResourcesFitArgs.ScoringStrategy=Topology) which models per-NUMA resources in the scheduler. - Use the scheduler-plugins/noderesourcetopology out-of-tree plugin which reads NodeResourceTopology CRs. - Or just use best-effort policy and live with occasional cross-NUMA placement.


14. ResourceQuota: Namespace Caps

ResourceQuota is an admission-controller mechanism that caps aggregate resource usage per namespace. It runs at the apiserver, before object creation. If a new pod would push the namespace over its quota, the create is rejected with 403 Forbidden.

14.1 What can be quota'd

apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-a-quota
  namespace: team-a
spec:
  hard:
    # Compute aggregate
    requests.cpu: "100"
    requests.memory: "200Gi"
    limits.cpu: "200"
    limits.memory: "400Gi"

    # Extended resources
    requests.nvidia.com/gpu: "8"

    # Storage aggregate
    requests.storage: "10Ti"
    persistentvolumeclaims: "50"

    # Object counts
    pods: "1000"
    services: "100"
    services.loadbalancers: "5"
    services.nodeports: "10"
    configmaps: "200"
    secrets: "200"
    replicationcontrollers: "100"

    # Per-StorageClass storage caps
    gold.storageclass.storage.k8s.io/requests.storage: "5Ti"
    silver.storageclass.storage.k8s.io/requests.storage: "10Ti"

    # Hugepages
    hugepages-2Mi: "10Gi"

    # Ephemeral storage
    requests.ephemeral-storage: "100Gi"
    limits.ephemeral-storage: "200Gi"

14.2 Two surprising rules

  • If a ResourceQuota exists for a resource (e.g., requests.cpu), every pod in that namespace MUST set that field. A pod missing requests.cpu is rejected at admission. This is by design: the apiserver needs to know how much quota to deduct. Combine with LimitRange (§15) to set defaults.
  • The quota is checked at create time, then a "used" counter is maintained. Subsequent updates that change requests trigger re-evaluation. If a pod's actual resource usage exceeds its requests, the quota is still happy — quota is about requests, not actual usage.

14.3 The accounting controller

The quotacontroller (pkg/controller/resourcequota/) watches pods and other objects, recomputes used quota, and writes it back to the ResourceQuota's status.used:

status:
  hard:
    requests.cpu: "100"
    requests.memory: 200Gi
    pods: "1000"
  used:
    requests.cpu: "47500m"
    requests.memory: 89Gi
    pods: "234"

kubectl describe quota team-a-quota shows this. When a pod creation is rejected:

Error from server (Forbidden): error when creating "deploy.yaml":
admission webhook "resourcequota.kubernetes.io" denied the request:
exceeded quota: team-a-quota,
requested: requests.cpu=8,
used: requests.cpu=98,
limited: requests.cpu=100

15. LimitRange: Per-Object Defaults and Bounds

LimitRange is the defaulting + per-object validating counterpart to ResourceQuota's aggregate role. It does three things:

  1. Sets default requests/limits for pods that omit them.
  2. Enforces min/max per container, per pod, per PVC.
  3. Enforces a maxLimitRequestRatio (limits ≤ N × requests).
apiVersion: v1
kind: LimitRange
metadata:
  name: team-a-limits
  namespace: team-a
spec:
  limits:
  - type: Container
    default:                  # used as limits if omitted
      cpu: "500m"
      memory: "512Mi"
    defaultRequest:           # used as requests if omitted
      cpu: "100m"
      memory: "128Mi"
    min:
      cpu: "10m"
      memory: "32Mi"
    max:
      cpu: "8"
      memory: "16Gi"
    maxLimitRequestRatio:
      cpu: "10"               # limit ≤ 10 × request
      memory: "4"             # limit ≤ 4 × request
  - type: Pod
    max:
      cpu: "16"               # aggregate across containers
      memory: "32Gi"
  - type: PersistentVolumeClaim
    min:
      storage: "1Gi"
    max:
      storage: "1Ti"

15.1 The defaulting interaction with QoS

A pod that specifies nothing in a namespace with a LimitRange:

# user submits this
apiVersion: v1
kind: Pod
metadata:
  name: bare
  namespace: team-a
spec:
  containers:
  - name: c
    image: nginx

…gets transformed by the LimitRange admission plugin into:

spec:
  containers:
  - name: c
    image: nginx
    resources:
      requests:    {cpu: "100m", memory: "128Mi"}    # from defaultRequest
      limits:      {cpu: "500m", memory: "512Mi"}    # from default

Now its QoS class is Burstable, not BestEffort. The user didn't ask for that; the LimitRange did it. This is the intent — LimitRange is how cluster admins make sure no pod in their namespace is BestEffort by accident.

15.2 Order: LimitRange runs before ResourceQuota

LimitRange is a mutating admission plugin (it modifies the pod). ResourceQuota is validating. Order matters: LimitRange fills in defaults first, then ResourceQuota deducts those defaults from quota. A pod with no resource spec gets the LimitRange defaults, those defaults are deducted from quota.


16. Quota Scopes and ScopeSelector

A bare ResourceQuota applies to every pod in the namespace. Sometimes you want quotas that apply only to specific kinds of pods. Hence scopes and scopeSelector.

16.1 The built-in scopes

Scope Matches pods that…
Terminating have spec.activeDeadlineSeconds >= 0
NotTerminating do not have activeDeadlineSeconds
BestEffort have QoS class BestEffort
NotBestEffort have QoS class Burstable or Guaranteed
PriorityClass have any non-empty priorityClassName
CrossNamespacePodAffinity use cross-namespace affinity

Example: cap BestEffort pod count without restricting compute:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: best-effort-pod-count
  namespace: team-a
spec:
  hard:
    pods: "100"
  scopes:
  - BestEffort

16.2 ScopeSelector with PriorityClass

apiVersion: v1
kind: ResourceQuota
metadata:
  name: high-priority-cpu
  namespace: team-a
spec:
  hard:
    requests.cpu: "200"     # team-a may use up to 200 CPU of high priority
  scopeSelector:
    matchExpressions:
    - scopeName: PriorityClass
      operator: In
      values: ["high"]

This is how you carve a namespace's allocation between priority classes. Combined with §17's PriorityClass-based preemption, you get a workable resource isolation between batch and serving workloads in the same namespace.

16.3 The Terminating gotcha

Pods with activeDeadlineSeconds set (typically Jobs, CronJob-owned pods) count against the Terminating scope. If you also have a separate quota matching NotTerminating, the same pod doesn't count against the latter. But the default quota (no scope) counts every pod. So:

  • Default quota: all pods → counted
  • Terminating quota: only Job-ish pods → counted
  • NotTerminating quota: only long-running pods → counted

A pod can count against multiple quotas at once.


17. PriorityClass and Preemption Recap

PriorityClass is a cluster-scoped object that maps a name to an integer priority. Pods reference it via spec.priorityClassName.

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: high
value: 1000000
globalDefault: false
description: "high-priority production workloads"
preemptionPolicy: PreemptLowerPriority

Defaults shipped with Kubernetes: - system-cluster-critical (2_000_000_000): essential cluster components. - system-node-critical (2_000_001_000): essential node components (kubelet, kube-proxy, CNI).

17.1 What priority does

  1. Scheduling order: the scheduler dequeues higher-priority pending pods first.
  2. Preemption: if a high-priority pod can't fit anywhere, the scheduler picks a node where evicting one or more lower-priority pods would make it fit. It evicts them (graceful delete, respecting PodDisruptionBudgets), waits, and then binds.
  3. Eviction immunity: the kubelet's eviction manager (§24) sorts victims by priority within QoS class. Higher-priority pods are evicted later.

17.2 Preemption respects PDBs

The scheduler runs a "would this PDB be violated?" check before nominating a victim. If preemption would put a Deployment below its minAvailable, the scheduler picks a different victim (or, if none exists, the preemption fails and the high-priority pod stays Pending).

17.3 The PreemptionPolicy

  • PreemptLowerPriority (default): high-priority pods preempt lower-priority pods to fit.
  • Never: the pod just queues. Won't kick anybody out. Used for high-priority batch (you want them ahead in the queue but you don't want them murdering services).

17.4 Without a PriorityClass, what's the priority?

If globalDefault: true is set on one PriorityClass, that's the default. Otherwise, default is 0. So existing pods (no priority) are at the bottom of the queue and the first to be preempted.


18. Extended and Scalar Resources

Not everything is CPU and memory. Extended resources are arbitrary integer-valued resources advertised by nodes and consumed by pods.

18.1 The advertisement side

A node controller, device plugin, or admin can write capacity for a custom resource:

# Manual patch (rare):
$ kubectl patch node node-1 --subresource=status --type='json' \
   -p='[{"op": "add", "path": "/status/capacity/example.com~1dongle", "value": "4"}]'

Real-world: device plugins do this automatically via kubelet's registration socket (/var/lib/kubelet/plugins_registry/). The GPU plugin advertises nvidia.com/gpu: 8 on each GPU node.

18.2 The consumption side

spec:
  containers:
  - name: train
    image: tensorflow/tensorflow:latest-gpu
    resources:
      requests:
        nvidia.com/gpu: 1
      limits:
        nvidia.com/gpu: 1     # required: requests must equal limits for extended

For extended resources, requests must equal limits. The scheduler treats them as binary "match or no-match" — fractional GPUs aren't allowed (you can't slice a GPU in two by saying 0.5 here; multi-tenancy on a GPU is the device plugin's job via MIG/MPS).

18.3 What the kubelet does

When the scheduler binds a pod requesting nvidia.com/gpu: 1, the kubelet:

  1. Calls the device plugin's Allocate(deviceIDs=[GPU-42]) gRPC.
  2. The plugin returns env vars (NVIDIA_VISIBLE_DEVICES=GPU-42), mount paths (/dev/nvidia0), and CDI annotations.
  3. The kubelet passes these into the CRI ContainerConfig.
  4. The runtime applies them when launching the container.

18.4 Scalar vs OCI device

"Extended resource" is the K8s API name. "Scalar resource" is sometimes used interchangeably. "OCI device" is the runtime concept — a device file mounted into the container. The Device Plugin maps K8s extended resources to OCI devices.


19. Hugepages: 2Mi and 1Gi

Linux normally manages memory in 4 KiB pages. For workloads with huge working sets (databases, JVMs, DPDK), 4 KiB pages cause TLB thrash — every access misses the TLB, costs a page-table walk, and burns CPU. Hugepages are 2 MiB or 1 GiB physical pages. Fewer TLB entries cover the same memory; databases see 10–20% throughput improvements.

19.1 Reservation at boot

Hugepages must be reserved from the kernel before any workload requests them. They are not allocatable from regular RAM on demand (without madvise(MADV_HUGEPAGE) and Transparent Huge Pages, which K8s doesn't use).

# In /etc/default/grub:
GRUB_CMDLINE_LINUX="default_hugepagesz=1G hugepagesz=1G hugepages=16 hugepagesz=2M hugepages=2048"
# Reserves: 16 × 1GiB pages and 2048 × 2MiB pages.

# Or at runtime (NUMA-aware):
echo 16 > /sys/devices/system/node/node0/hugepages/hugepages-1048576kB/nr_hugepages
echo 16 > /sys/devices/system/node/node1/hugepages/hugepages-1048576kB/nr_hugepages

19.2 Advertised by the kubelet

After boot, the kubelet auto-discovers hugepages and adds them to node.status.capacity:

status:
  capacity:
    cpu: "32"
    memory: "131072Mi"
    hugepages-1Gi: "16Gi"
    hugepages-2Mi: "4Gi"

19.3 Consumed by pods

spec:
  containers:
  - name: db
    image: postgres
    resources:
      requests:
        memory: "10Gi"
        hugepages-1Gi: "8Gi"     # 8 × 1GiB pages
      limits:
        memory: "10Gi"
        hugepages-1Gi: "8Gi"
    volumeMounts:
    - mountPath: /hugepages
      name: hugepage-1gi
  volumes:
  - name: hugepage-1gi
    emptyDir:
      medium: HugePages-1Gi      # the file-backed view of hugepages

The pod gets a tmpfs-style mount backed by hugepages; the application opens a file there and mmap()s it. Pods using hugepages must be Burstable or Guaranteed (BestEffort can't request them).

19.4 Hugepages and QoS

Hugepages don't participate in QoS class derivation (§3). A pod with hugepages but no CPU/memory requests is BestEffort and can be evicted; the hugepages are released back to the pool on eviction. Treat hugepages as a capacity-bound resource, not a QoS one.


20. Ephemeral Storage

Every container has a writable layer (the OCI runtime's overlay upperdir) and (usually) some emptyDir volumes. Both live on the node's filesystem and count against ephemeral storage.

20.1 What counts

  • Container writable layer (anything written outside a mounted volume).
  • emptyDir volumes (unless medium: Memory, which uses tmpfs and counts against memory).
  • Container logs (/var/log/pods/<pod>/<container>/0.log).

What does not count: persistent volumes (those are CSI's problem), hostPath volumes, emptyDir with tmpfs medium.

20.2 Requests and limits

resources:
  requests:
    ephemeral-storage: "1Gi"
  limits:
    ephemeral-storage: "5Gi"

Scheduler: counts requests.ephemeral-storage against node.status.allocatable.ephemeral-storage.

Kubelet: every 10s the eviction manager runs du/statfs on the container's writable layer and emptyDirs. If usage exceeds limits.ephemeral-storage, the kubelet evicts the pod (sends kubectl delete --grace-period=...). This is a per-pod eviction, not a node-level one.

20.3 The cost: du is expensive

The kubelet walking the pod's writable layer with du is O(files). At 100 pods × 50k files each = 5M stat() calls every 10s. On slow disks this is noticeable. Newer kubelets (1.31+) use filesystem-level quotas (project quotas on xfs/ext4) when available, which give O(1) accounting. Enable via featureGates: LocalStorageCapacityIsolationFSQuotaMonitoring: true.

20.4 nodefs vs imagefs

The node has two storage partitions in the kubelet's mental model: - nodefs: where pod root volumes, emptyDirs, logs live. Usually /var/lib/kubelet. - imagefs: where the container runtime stores images and writable layers. Usually /var/lib/containerd or /var/lib/docker.

If both are on the same filesystem (common), they're the same partition. If split (recommended for production), imagefs.available and nodefs.available are tracked separately — see §24.


21. PID Limits and pid.available

Linux's global PID space is bounded (/proc/sys/kernel/pid_max, default 4194304 on modern kernels). Each cgroup-v2 also has pids.max. Exhaust either and fork() returns EAGAIN. New processes can't start; the kubelet itself may fail to launch new containers.

21.1 Per-pod limit: podPidsLimit

# KubeletConfiguration
podPidsLimit: 4096

Default: -1 (unlimited inside the cgroup; only pid_max caps you). With this set, every container's cgroup gets pids.max = 4096. If a pod tries to fork the 4097th task, EAGAIN.

$ cat /sys/fs/cgroup/kubepods.slice/.../pids.max
4096
$ cat /sys/fs/cgroup/kubepods.slice/.../pids.current
237

21.2 Node-level: pid.available eviction signal

The kubelet tracks total node PIDs vs pid_max. When the ratio crosses an eviction threshold, pods get evicted. Default thresholds:

evictionHard:
  pid.available: "10%"        # default: 10% of pid_max remaining
evictionSoft:
  pid.available: "15%"
evictionSoftGracePeriod:
  pid.available: "1m"

21.3 Common PID exhaustion patterns

  • Java apps with high thread counts. Each thread is a task. 1000 threads × 100 pods = 100k. Within default limits, fine. With podPidsLimit: 1024, every other pod will hit it during GC.
  • Process forking shell pipelines. Old-school shell scripts that fork hundreds of subshells per request.
  • Misconfigured workers. Gunicorn/Unicorn with thousands of workers.
  • Zombie process accumulation. A container that fork()s but never wait()s leaves zombies. They count against pids.current until the parent dies.

22. The Kubelet's Reservation Model: Allocatable

Not every byte of node RAM is available to pods. The kubelet reserves chunks for itself, for the OS, and for an eviction buffer.

node.status.capacity         = total physical resources
node.status.allocatable      = capacity − kubeReserved − systemReserved − evictionHard

22.1 The KubeletConfiguration knobs

kubeReserved:                  # for kubelet, container runtime, plugins
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "10Gi"
  pid: "1000"
systemReserved:                # for the OS (systemd, sshd, kernel...)
  cpu: "500m"
  memory: "1Gi"
  ephemeral-storage: "10Gi"
  pid: "1000"
evictionHard:                  # hard floor (eviction triggers immediately if crossed)
  memory.available: "500Mi"
  nodefs.available: "10%"
  nodefs.inodesFree: "5%"
  imagefs.available: "15%"
  pid.available: "10%"

Now allocatable = capacity − (500m+500m) − (1Gi+1Gi) − 500Mi (and similar for storage/PID).

22.2 The cgroups they live in

The kubelet creates a cgroup hierarchy that enforces these reservations:

/sys/fs/cgroup/
├── kubepods.slice/             ← cpu.weight set so pods get (capacity - kubeReserved - systemReserved)
├── system.slice/               ← OS workloads
│   ├── kubelet.service/        ← reserved by KubeReserved (via systemd Slice= or kubeletCgroups)
│   └── containerd.service/
└── ...

Two enforcement modes (enforceNodeAllocatable): - ["pods"] (default): only the kubepods.slice cgroup is constrained. The kubelet/system can still grow. - ["pods", "kube-reserved", "system-reserved"]: also constrain the kubelet's and system's cgroups. Stronger isolation but risks killing the kubelet under load — used carefully.

enforceNodeAllocatable: ["pods"]
kubeletCgroups: "/kubelet.slice"
systemCgroups: "/system.slice"

22.3 The math

On a 32-core, 128 GiB node with the config above:

Capacity:    cpu=32, memory=128Gi, ephemeral-storage=400Gi, pid=4194304
Reserved:    cpu=1, memory=2Gi, ephemeral-storage=20Gi, pid=2000
Eviction:    memory=500Mi, ephemeral-storage=40Gi (10% nodefs), pid=419k (10% pid)

Allocatable: cpu=31, memory=125.5Gi, ephemeral-storage=340Gi, pid=3.77M

The scheduler sees allocatable. The pods can use at most that. The reserved 2.5Gi + 500Mi = 3Gi of memory is invisible to pods but is what keeps the kubelet alive when pods misbehave.


23. Cgroup Hierarchy on a Kubernetes Node

We already showed the pod-level hierarchy in §6; here's the full picture including system reservations:

/sys/fs/cgroup/                                       ← root cgroup-v2
│   cpu.max = max
│   memory.max = max
│   memory.current = 23 GiB                            ← total node memory used
│
├── init.scope/                                       ← PID 1
│
├── system.slice/                                     ← OS + agent processes
│   │   cpu.weight = 100 (default — could be raised by enforceNodeAllocatable)
│   │   memory.max = (systemReserved.memory) if enforced
│   │
│   ├── kubelet.service/                              ← the kubelet process
│   │   │   cpu.weight = 100
│   │   │   memory.current = 220 MiB
│   │   │
│   │   └── tasks
│   │
│   ├── containerd.service/
│   ├── sshd.service/
│   ├── systemd-resolved.service/
│   └── ...
│
├── user.slice/                                       ← interactive logins
│
└── kubepods.slice/                                   ← all pods
    │   cpu.weight = 1000 - 100 (system) - 100 (kube)
    │   memory.max = max OR (allocatable.memory) if enforced
    │   memory.current = sum of pod memory.current
    │
    ├── kubepods-besteffort.slice/                    ← BestEffort QoS group
    │   │   cpu.weight = 1
    │   │   memory.max = max
    │   │
    │   ├── kubepods-besteffort-pod<UID>.slice/
    │   └── ...
    │
    ├── kubepods-burstable.slice/                     ← Burstable QoS group
    │   │   cpu.weight = computed (between BestEffort and Guaranteed)
    │   │   memory.max = max
    │   │
    │   ├── kubepods-burstable-pod<UID>.slice/
    │   │   ├── cri-containerd-<pause>.scope/
    │   │   ├── cri-containerd-<init>.scope/
    │   │   └── cri-containerd-<app>.scope/
    │   │           cpu.weight = (request.cpu derived)
    │   │           cpu.max    = (limit.cpu derived) or "max"
    │   │           memory.max = (limit.memory)
    │   │           memory.high = (with MemoryQoS)
    │   │           memory.min  = (with MemoryQoS)
    │   │           pids.max   = podPidsLimit
    │   │           cpuset.cpus = (with static CPU manager: shared pool)
    │   │
    │   └── ...
    │
    └── kubepods-pod<UID>.slice/                      ← Guaranteed pod (direct child!)
        │   cpu.weight = sum of requests
        │   memory.max = sum of limits
        │   cpuset.cpus = 4-7 (pinned)
        │   cpuset.mems = 0   (NUMA pinned)
        │
        └── cri-containerd-<app>.scope/
                cpu.weight = ...
                cpu.max    = ...
                memory.max = ...

23.1 Pid namespace vs cgroup namespace

Note that the cgroup hierarchy is orthogonal to the PID namespace. A container has its own PID namespace (PID 1 inside, mapped to some host PID outside). The cgroup namespace controls which slice of the cgroup tree the container sees in its own /proc/self/cgroup. The kubelet's containers have cgroup namespaces enabled by default (since 1.20).


24. Eviction Signals and Thresholds

The kubelet's eviction manager (pkg/kubelet/eviction/) periodically polls node-level resource signals and, if any threshold is crossed, evicts pods until pressure resolves. This is the proactive counterpart to the kernel's reactive OOM killer.

24.1 The six core signals

Signal Source What it means
memory.available cgroupfs (root cgroup) bytes of memory free at node level
nodefs.available statfs on root filesystem free bytes on /var/lib/kubelet
nodefs.inodesFree statfs on root filesystem free inodes on /var/lib/kubelet
imagefs.available statfs on imagefs (if split) free bytes on /var/lib/containerd
imagefs.inodesFree statfs on imagefs free inodes
pid.available /proc/sys/kernel/pid_max unused PIDs
allocatableMemory.available aggregated cgroup pod stats distinct from node memory.available

24.2 Hard vs soft thresholds

evictionHard:                  # cross → immediate eviction (no grace period)
  memory.available: "100Mi"
  nodefs.available: "10%"
  nodefs.inodesFree: "5%"
  imagefs.available: "15%"
  imagefs.inodesFree: "5%"
  pid.available: "10%"

evictionSoft:                  # cross → eviction after evictionSoftGracePeriod
  memory.available: "300Mi"
  nodefs.available: "15%"
  pid.available: "15%"

evictionSoftGracePeriod:       # how long the signal must be over threshold
  memory.available: "1m30s"
  nodefs.available: "2m"
  pid.available: "1m30s"

evictionMaxPodGracePeriod: 60   # how long evicted pods get to terminate
evictionPressureTransitionPeriod: 5m   # cooldown after pressure resolved

Hard eviction means the kubelet uses gracePeriodOverride=0 — pods get SIGKILL immediately. Soft eviction means pods get their normal terminationGracePeriodSeconds, capped at evictionMaxPodGracePeriod.

24.3 The decision tree

                  ┌──────────────────────────────────┐
                  │  Eviction manager runs every 10s │
                  └────────────────┬─────────────────┘
                                   │
              ┌────────────────────┴────────────────────┐
              ▼                                         ▼
   ┌────────────────────┐                  ┌───────────────────────┐
   │  Read signals:     │                  │  Read pod usage from  │
   │  - memory.available│                  │  cgroup stats + statfs│
   │  - nodefs.*        │                  │  per pod              │
   │  - imagefs.*       │                  └───────────┬───────────┘
   │  - pid.available   │                              │
   └────────┬───────────┘                              │
            │                                          │
            ▼                                          │
   ┌────────────────────────────────────────────────┐  │
   │  For each signal:                              │  │
   │    if signal < evictionHard[signal] → HARD     │  │
   │    elif signal < evictionSoft[signal] for      │  │
   │         > evictionSoftGracePeriod → SOFT       │  │
   └────────┬───────────────────────────────────────┘  │
            │                                          │
            ▼                                          ▼
   ┌────────────────────────────────────────────────────────────┐
   │  rank pods by eviction priority (§25):                     │
   │  1. Resource type matches signal (e.g., memory pressure → │
   │     evict by memory usage; disk pressure → by disk usage) │
   │  2. QoS class: BestEffort < Burstable < Guaranteed        │
   │  3. Priority (lower priority dies first)                  │
   │  4. Usage above requests (more overshoot dies first)      │
   │  5. Pod start time (older dies first, last)               │
   └────────┬───────────────────────────────────────────────────┘
            │
            ▼
   ┌─────────────────────────────────────────┐
   │  Pick top victim.                       │
   │  Evict via apiserver delete with        │
   │  appropriate grace period.              │
   │  Update node condition:                 │
   │    MemoryPressure | DiskPressure |      │
   │    PIDPressure                          │
   │  Re-evaluate next tick.                 │
   └─────────────────────────────────────────┘

24.4 Node conditions during pressure

When the eviction manager detects pressure, it sets node conditions:

status:
  conditions:
  - type: MemoryPressure
    status: "True"
    lastHeartbeatTime: "2025-05-23T14:32:01Z"
    reason: KubeletHasInsufficientMemory
  - type: DiskPressure
    status: "False"
  - type: PIDPressure
    status: "False"

The scheduler watches these conditions and avoids placing new pods on nodes with MemoryPressure=True (BestEffort) or DiskPressure=True (everything). The kubelet rejects new pods at admission while pressure persists.

24.5 The cooldown

evictionPressureTransitionPeriod (default 5m) prevents flapping: once pressure resolves, the kubelet waits 5 minutes before clearing the node condition. Otherwise a node oscillating around the threshold would oscillate the condition, the scheduler would oscillate placements, and so on.


25. Eviction Ranking: BestEffort First

The eviction manager's ranking algorithm (pkg/kubelet/eviction/helpers.go, rankMemoryPressure etc.) is roughly:

// Pseudocode of eviction ranking under memory pressure
func rankPodsForMemoryEviction(pods []*v1.Pod) []*v1.Pod {
    sort.SliceStable(pods, func(i, j int) bool {
        // 1. Critical pods never evicted.
        if isCritical(pods[i]) { return false }
        if isCritical(pods[j]) { return true }

        // 2. QoS class: BestEffort < Burstable < Guaranteed.
        qi, qj := qos(pods[i]), qos(pods[j])
        if qi != qj {
            return qosOrder(qi) < qosOrder(qj)   // BestEffort first
        }

        // 3. Pod priority (lower priority = die first).
        pi, pj := priority(pods[i]), priority(pods[j])
        if pi != pj {
            return pi < pj
        }

        // 4. Memory usage relative to request.
        //    More overshoot of request = die first.
        oi := memUsage(pods[i]) - memRequest(pods[i])
        oj := memUsage(pods[j]) - memRequest(pods[j])
        if oi != oj {
            return oi > oj
        }

        // 5. Tie-break by pod start time (older survives longer? actually
        //    older first to die — but in practice this rarely tips the scale).
        return pods[i].Status.StartTime.Before(pods[j].Status.StartTime)
    })
    return pods
}

For disk pressure, the ranking uses ephemeral-storage usage instead of memory.

For PID pressure, by process count.

25.1 Why "overshoot of request"?

Two Burstable pods, same priority, on the same node, both contributing to memory pressure:

  • Pod A: requests.memory=1Gi, limits.memory=2Gi, currently using 1.5 GiB.
  • Pod B: requests.memory=1Gi, limits.memory=2Gi, currently using 1.9 GiB.

A's overshoot = 0.5 GiB. B's overshoot = 0.9 GiB. B dies first. The logic: B is more responsible for the pressure because it asked for the same amount but is consuming more.

25.2 Guaranteed pods are not invulnerable

A Guaranteed pod is last in the eviction queue, but it can still be evicted if: - The node has only Guaranteed pods and is still under pressure. - Or the Guaranteed pod itself is using more memory than its limit (which is impossible — the kernel would have OOM-killed it first — but disk usage can exceed limit briefly).

Generally a node under pressure with only Guaranteed pods means your scheduler over-promised; this is rare and indicates a configuration bug (e.g., kubeReserved too low).

25.3 Eviction events

$ kubectl get events --field-selector reason=Evicted
LAST SEEN   TYPE     REASON   OBJECT          MESSAGE
3m12s       Warning  Evicted  pod/foo-abc     The node was low on resource: memory. Container foo was using 1.2Gi, which exceeds its request of 512Mi.

These events live in the namespace of the evicted pod. Always shipped to long-term storage; they vanish from etcd after 1 hour by default.


26. Eviction vs OOM Kill: Proactive vs Reactive

Two mechanisms exist for memory pressure. They look similar but they're different:

Property Eviction OOM kill
Triggered by kubelet polling kernel detecting pressure
Granularity whole pod one process in a cgroup
Notice grace period (or zero) immediate SIGKILL
Selection QoS class + priority + over-request oom_score_adj
Pod restartPolicy respected (pod terminates) respected (container restart)
Status reason Evicted OOMKilled (container) or kernel msg
Triggering threshold evictionHard.memory.available (e.g., 500Mi free at node) memory.max exceeded for a cgroup

The relationship:

   ↑ memory pressure
   │                            ┌─────────────────────────┐
                                │  cgroup OOM kill        │
                                │  (a pod exceeded its    │
                                │   memory.max — its own  │
                                │   pod-local cgroup)     │
                                └─────────────────────────┘

   ────  evictionHard.memory.available threshold (e.g., 500 MiB) ────

                                ┌─────────────────────────┐
                                │  Kubelet eviction       │
                                │  (node-wide pressure;   │
                                │   kubelet picks victim) │
                                └─────────────────────────┘

   ────  evictionSoft.memory.available threshold (e.g., 1 GiB) ────

                                graceful eviction with delay

   ↓ memory headroom

The cgroup OOM is per-pod and happens whenever an individual pod exceeds its own memory.max. It doesn't need node-level pressure. A single pod with limits.memory: 100Mi will get OOM-killed when it tries to use 101 MiB, even if the node has 100 GiB of free RAM.

Kubelet eviction is node-wide. It only activates when aggregate memory pressure is high. It evicts pods preemptively so the kernel OOM never has to fire — the kernel OOM is a last resort that might kill the wrong thing.

26.1 The order in a real cascade

  1. Node has 64 GiB; 60 GiB used by pods; 4 GiB free.
  2. A Burstable pod with limits.memory: 8Gi starts allocating fast.
  3. As that pod's cgroup approaches its 8 GiB limit, the cgroup OOM may fire before the node-level signal moves much. Pod dies. Memory freed.
  4. Or: pod is well-behaved, only uses 6 GiB, but two other pods grow simultaneously. Node free drops below evictionHard.memory.available=500Mi.
  5. Kubelet eviction manager fires: picks the worst BestEffort or worst-overshoot Burstable. Evicts. Free memory recovers.
  6. Or: eviction is too slow (allocator burns memory at 5 GiB/s). Kernel global OOM fires before kubelet can act. The kernel picks the highest oom_score_adj (BestEffort or big-Burstable). That process dies.

The whole design is layered so that the kernel OOM is only invoked when both the kubelet eviction and the cgroup OOM failed to catch it.


27. Throttling Timeline and cpu.stat

A concrete trace of CPU throttling, end-to-end, observable from the metrics.

Pod spec:
  resources:
    requests: { cpu: 100m }
    limits:   { cpu: 400m }

cpu.max = "40000 100000"   (40ms quota / 100ms period)
cpu.weight = 4              (from 100m request)

Workload: 4-thread Java app, each thread does ~30ms of CPU work per request.

Period 1: [00ms-100ms]
   t=0    request arrives, 4 threads spin up
   t=0    all 4 threads execute in parallel
   t=10ms each thread has done 10ms wall × 4 = 40ms cgroup CPU; QUOTA HIT
   t=10ms kernel throttles ALL tasks in cgroup until end of period
   t=100ms next period begins; cumulative wall time: 100ms
          per-thread CPU used: 10ms × 4 = 40ms of 120ms requested
          request still not complete

Period 2: [100ms-200ms]
   t=100ms threads resume, do another 10ms each
   t=110ms quota hit again
   t=200ms next period
          per-thread CPU used: 20ms each; still 30ms requested

Period 3:
   t=200ms threads resume, do another 10ms
   t=210ms quota hit
   t=300ms next period

Period 4:
   t=300ms threads resume, do final 0ms
   request complete

WALL CLOCK: 300ms for a request that needed 30ms of CPU.
THROTTLED:  3 periods × 90ms = 270ms of throttle.
PERIODS:    3 throttled out of 3 = 100% throttling rate.

cpu.stat after this request:
   nr_periods       3 (more, this is just our delta)
   nr_throttled     3
   throttled_usec   270000

The user-visible symptom: P99 latency 300ms for a workload that benchmarks at 30ms in isolation. Root cause: 4 threads × limits.cpu=400m = 100m per thread effective; threads complete only 25% as fast as their isolated benchmark.

27.1 The metric in Prometheus

container_cpu_cfs_throttled_periods_total{container="app",pod="foo"} 12847
container_cpu_cfs_periods_total{container="app",pod="foo"} 14223

Throttle rate = 12847 / 14223 = 90.3%.

27.2 The PSI signal

$ cat /sys/fs/cgroup/kubepods.slice/.../cpu.pressure
some avg10=78.20 avg60=72.14 avg300=68.32 total=482310291
full avg10=2.10 avg60=1.42 avg300=1.18 total=8429101

some avg10=78% means: over the last 10 seconds, the cgroup had at least one task waiting for CPU 78% of the time. That's a severe indicator of CFS throttling or CPU contention.

27.3 Prometheus alert

- alert: ContainerCPUThrottled
  expr: |
    rate(container_cpu_cfs_throttled_periods_total{container!="",container!="POD",image!=""}[5m]) /
    rate(container_cpu_cfs_periods_total{container!="",container!="POD",image!=""}[5m]) > 0.25
  for: 15m
  labels:
    severity: warning
  annotations:
    summary: "Container {{ $labels.namespace }}/{{ $labels.pod }}/{{ $labels.container }} CFS-throttled"
    description: |
      More than 25% of CFS periods ended in throttle for 15 minutes.
      This usually indicates limits.cpu is too low, or that the workload
      is multi-threaded and exceeds limits in burst.
      Action: either raise limits.cpu, remove the limit, or investigate
      the workload's burst pattern (GC pauses, async I/O storms, etc).

28. In-Place Pod Resize

Historically, changing a pod's resources required deleting and recreating the pod. Since 1.27 (alpha) and 1.32 (GA), Kubernetes supports in-place resize: change resources on a running pod and have the kubelet adjust cgroup files without restarting.

28.1 The resize subresource

$ kubectl patch pod my-pod --subresource resize --patch '
spec:
  containers:
  - name: app
    resources:
      requests: { cpu: "500m", memory: "1Gi" }
      limits:   { cpu: "1",    memory: "2Gi" }
'

This goes to a special apiserver subresource (/api/v1/.../pods/<name>/resize) that bypasses the usual immutability of pod.spec.containers[].resources.

28.2 Resize policy

The pod can declare per-resource resize behavior:

spec:
  containers:
  - name: app
    resources:
      requests: { cpu: "100m", memory: "256Mi" }
      limits:   { cpu: "500m", memory: "512Mi" }
    resizePolicy:
    - resourceName: cpu
      restartPolicy: NotRequired       # default: in-place
    - resourceName: memory
      restartPolicy: RestartContainer  # required if memory needs restart

Memory resize is sometimes risky (JVMs, RocksDB, malloc arenas may not handle dynamic shrink) — the operator can say "for memory changes, restart the container."

28.3 What the kubelet does

When the apiserver applies the patch, the kubelet's pod worker:

  1. Recomputes the cgroup values.
  2. Checks if the node has capacity for the new request (admits or denies).
  3. If admitted: writes new cpu.weight, cpu.max, memory.max to the cgroup files.
  4. Updates pod.status.containerStatuses[].resources to reflect the actual applied values.

For memory shrinking: the kubelet attempts to write the new (smaller) memory.max. If memory.current > new memory.max, the kernel forces reclaim and may OOM-kill. The kubelet can be configured to detect this and refuse the resize.

28.4 Resize status conditions

The pod gets two new conditions during a resize:

status:
  conditions:
  - type: PodResizePending
    status: "True"
    reason: Deferred
    message: "Node cannot fit new size now"
  - type: PodResizeInProgress
    status: "True"

VPA (§29) is the primary consumer of this API. Without VPA, manual in-place resize is rarely useful in production.


29. VPA Integration (Forward Ref)

The Vertical Pod Autoscaler (autoscaling.k8s.io/v1) is the primary consumer of in-place resize. It has three components:

  • Recommender: watches pod metrics, computes recommended requests/limits from a histogram of past usage (default: 90th percentile + safety margin).
  • Updater: evicts pods whose current requests differ significantly from recommendations.
  • Admission Controller: rewrites resource requests on pod create, using the recommender's output.

With in-place resize (1.32+), VPA Updater can resize in-place instead of evicting — much smoother, no downtime.

Chapter 22 covers VPA in full. The key interaction with this chapter: VPA reads the metrics we discuss in §31, applies the right-sizing heuristic from §30, and writes back into the pod's resources block.

A non-obvious caveat: VPA conflicts with HPA on the same metric. HPA scales replicas based on CPU usage; VPA changes requests, which changes the base of the HPA's "% of request" metric. Result: oscillation. Mitigation: VPA in Off (recommend-only) mode + HPA controlling replicas; or VPA on memory + HPA on CPU.


30. The Right-Sizing Workflow

How do you actually decide what numbers to put in your spec? The workflow that works:

30.1 Measure

Deploy with generous requests and no CPU limits (or a generous one) for at least a week. Capture:

  • p50, p95, p99 CPU usage (from container_cpu_usage_seconds_total).
  • p50, p95, p99 memory working set (from container_memory_working_set_bytes).
  • Throttling rate (from container_cpu_cfs_throttled_periods_total).
# p99 CPU usage over 7 days
quantile_over_time(0.99,
  rate(container_cpu_usage_seconds_total{namespace="prod",pod=~"my-app-.*"}[5m])[7d:5m])

# p99 memory working set
quantile_over_time(0.99,
  container_memory_working_set_bytes{namespace="prod",pod=~"my-app-.*"}[7d])

30.2 Set requests

  • requests.cpu ≈ p99 of measured CPU usage (or p95, depending on how aggressive you want bin-packing). This guarantees scheduler-promised CPU = peak-needed CPU.
  • requests.memory ≈ p99 of working-set memory, rounded up.

For predictable workloads, requests = p99 means you essentially never hit contention. For bursty workloads, you may want requests = p50 and rely on cpu.weight + headroom to absorb bursts.

30.3 Set limits

  • limits.cpu: omit for latency-sensitive services. Set to 2 × requests for cap-able batch.
  • limits.memory ≈ 1.5 × requests.memory. The 50% headroom is your safety margin: it accommodates allocator overhead, GC peaks, and occasional anomalies, while still bounding worst-case node impact.

30.4 Iterate

After deploying with these numbers, re-measure for a week. Watch: - container_oom_events_total (memory limit too tight). - kube_pod_container_status_terminated_reason{reason="OOMKilled"} (same, with backoff). - container_cpu_cfs_throttled_periods_total (CPU limit too tight). - Evictions (kubelet_evictions{eviction_signal=...}).

If any of those fire, adjust.

30.5 The pattern in YAML

spec:
  containers:
  - name: api
    image: my/api:1.2.3
    resources:
      requests:
        cpu: "500m"        # ~p95 measured usage
        memory: "1Gi"      # ~p99 measured working set
      limits:
        # no cpu limit (latency-sensitive)
        memory: "1500Mi"   # 1.5× request

For batch:

spec:
  containers:
  - name: trainer
    image: my/trainer:1.0
    resources:
      requests:
        cpu: "4"
        memory: "16Gi"
      limits:
        cpu: "8"           # cap; batch can be throttled
        memory: "24Gi"

31. Observability: Metrics, Alerts, Dashboards

31.1 The essential metrics

From kube-state-metrics (object-level): - kube_pod_container_resource_requests{resource="cpu"} - kube_pod_container_resource_requests{resource="memory"} - kube_pod_container_resource_limits{resource="cpu"} - kube_pod_container_resource_limits{resource="memory"} - kube_pod_status_qos_class - kube_resourcequota{resource="...",type="used"} and type="hard"

From cAdvisor (/metrics/cadvisor) (actual usage): - container_cpu_usage_seconds_total (counter) - container_memory_working_set_bytes (gauge, the "RSS that counts") - container_memory_rss (gauge, the "RSS without page-cache") - container_memory_cache (gauge) - container_memory_max_usage_bytes (gauge, the high-watermark) - container_cpu_cfs_throttled_periods_total - container_cpu_cfs_throttled_seconds_total - container_cpu_cfs_periods_total - container_oom_events_total - container_fs_usage_bytes - container_fs_inodes_free

From kubelet (/metrics): - kubelet_evictions{eviction_signal="memory.available"} - kubelet_node_name (target labeling) - kubelet_running_pods - kubelet_running_containers

From the kernel / node-exporter: - node_memory_MemAvailable_bytes - node_filesystem_avail_bytes - node_filesystem_inodes_free - node_filesystem_files_free

31.2 Essential alerts

groups:
- name: pod-resources
  rules:

  - alert: PodMemoryUsageNearLimit
    expr: |
      (container_memory_working_set_bytes{container!="",container!="POD"}
       /
       on(pod, container, namespace) kube_pod_container_resource_limits{resource="memory"}) > 0.9
    for: 10m
    labels: { severity: warning }
    annotations:
      summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} near memory limit"
      description: |
        Container '{{ $labels.container }}' is using >90% of its memory limit
        for 10 minutes. OOMKill is likely. Investigate or raise limit.

  - alert: PodOOMKilled
    expr: |
      increase(container_oom_events_total[5m]) > 0
    labels: { severity: critical }
    annotations:
      summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} OOMKilled"
      description: "OOM event detected; check memory.max sizing."

  - alert: PodCPUThrottlingHigh
    expr: |
      (rate(container_cpu_cfs_throttled_periods_total[5m])
       /
       rate(container_cpu_cfs_periods_total[5m])) > 0.25
    for: 15m
    labels: { severity: warning }
    annotations:
      summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} CPU-throttled"
      description: |
        Container '{{ $labels.container }}' has been throttled >25% of periods
        for 15m. Investigate limit sizing or remove limit for this workload.

  - alert: NodeMemoryPressure
    expr: kube_node_status_condition{condition="MemoryPressure",status="true"} == 1
    for: 5m
    labels: { severity: warning }
    annotations:
      summary: "Node {{ $labels.node }} under memory pressure"
      description: "Eviction manager has marked the node MemoryPressure=True."

  - alert: NodeDiskPressure
    expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
    for: 5m
    labels: { severity: warning }
    annotations:
      summary: "Node {{ $labels.node }} under disk pressure"

  - alert: NodePIDPressure
    expr: kube_node_status_condition{condition="PIDPressure",status="true"} == 1
    for: 5m
    labels: { severity: warning }
    annotations:
      summary: "Node {{ $labels.node }} under PID pressure"

  - alert: ResourceQuotaNearLimit
    expr: |
      (kube_resourcequota{type="used"} / on(namespace,resource,resourcequota)
       kube_resourcequota{type="hard"}) > 0.9
    for: 30m
    labels: { severity: warning }
    annotations:
      summary: "ResourceQuota {{ $labels.namespace }}/{{ $labels.resourcequota }} near limit"
      description: "{{ $labels.resource }} usage > 90% of quota."

  - alert: PodEvicted
    expr: |
      increase(kubelet_evictions{eviction_signal!=""}[10m]) > 0
    labels: { severity: warning }
    annotations:
      summary: "Pods evicted on node {{ $labels.node }} (signal: {{ $labels.eviction_signal }})"
      description: "Investigate which workloads are over-promised."

  - alert: NodeAllocatableExhausted
    expr: |
      sum by (node) (kube_pod_container_resource_requests{resource="memory"})
      /
      kube_node_status_allocatable{resource="memory"} > 0.95
    for: 15m
    labels: { severity: warning }
    annotations:
      summary: "Node {{ $labels.node }} memory >95% promised"
      description: "Requests exhausting allocatable; pods may not schedule."

31.3 The "requests vs usage" right-sizing dashboard

A single dashboard with two side-by-side panels per workload:

  • Panel 1: kube_pod_container_resource_requests{resource="cpu"} plotted against rate(container_cpu_usage_seconds_total[5m]). Big gap (request >> usage) = waste. No gap (usage tracking request) = sized correctly. Usage > request = under-promised; either raise request or expect throttling.
  • Panel 2: same for memory.

VPA recommendations can be plotted alongside as a third line. This dashboard pays for itself in cluster compute savings within a month.


32. Pitfalls

The 25+ ways resource management goes wrong in production. Some of these come up once a year; some come up every week.

32.1 No requests — every node looks "free"

The scheduler treats unspecified requests.cpu/requests.memory as 0. Workload runs, uses real CPU and memory, but the node's allocatable - requested stays at "almost full" because nothing is reserved. More pods get scheduled. Node hits eviction territory. Pods get evicted (BestEffort first). Pager fires at 3am. → Always set requests. Use LimitRange to provide defaults.

32.2 CPU limit causing tail latency

The single most-common subtle bug. Multi-threaded app + limits.cpu set tight + occasional bursts = CFS throttles entire app for the rest of the 100ms period, blowing tail latency. → Remove CPU limits on latency-sensitive workloads. Memory limits stay.

32.3 Memory limit at working-set size — guaranteed OOM

Setting limits.memory = avg_working_set means the first GC pause / cache spike / connection storm causes an OOM. → Limit ≈ 1.5× p99 working set. Or: don't set a limit and rely on eviction (but then a single bad pod can take down the node).

32.4 QoS BestEffort in production

A team copy-pastes a Helm chart without resource blocks. Pod is BestEffort. Eviction manager picks it as the first victim under any pressure. Service goes into CrashLoopBackOff. → Use LimitRange to default to Burstable.

32.5 Missing ephemeral-storage requests

Pod writes 50 GiB of logs into /var/log (which is in the writable layer). Node's nodefs.available drops. Eviction fires, kills other pods. → Set requests.ephemeral-storage and limits.ephemeral-storage; ship logs externally.

32.6 Low podPidsLimit

podPidsLimit: 1024 looked safe until a Java app spawned 800 threads. pids.max hit; new thread → fork() returns EAGAIN; app crashes in non-obvious ways (typically NIO selector init failure). → Default to 4096, raise for heavy-thread apps.

32.7 cgroup-v1 vs v2 differences

Old kernels use cgroup-v1. memory.high is v2-only. CFS quota semantics are subtly different (v1: cpu.cfs_quota_us and cpu.cfs_period_us; v2: cpu.max is one file). Tools (kubectl-tree, kubectl-resource, custom scripts) that hardcode v1 paths break on v2. → Standardize on v2; require kernel ≥5.8.

32.8 LimitRange too strict — pod creation fails

LimitRange with max.cpu: 1 prevents any legitimate compute-heavy pod from being created in the namespace. Error: pods "x" is forbidden: maximum cpu usage per Container is 1, but limit is 4. → LimitRange is a guardrail, not a budget; size it generously, use ResourceQuota for actual budgets.

32.9 ResourceQuota counts Terminating pods

A Job creates a pod that's Completed but not yet garbage-collected; it still counts against requests.cpu in the quota. New job submission fails. → Use ttlSecondsAfterFinished on Jobs, or a quota scope NotTerminating to exclude them.

32.10 Object-count quota exhaustion on PVC churn

persistentvolumeclaims: 50 quota. CI creates 50 PVCs/hour with ttlSecondsAfterFinished: 0 Jobs. PVCs not GC'd in time. Quota saturates. New CI runs fail. → Lower TTL, raise quota, or use ephemeral PVCs.

32.11 PriorityClass without resource budget

Setting high priority on a Deployment means the scheduler will preempt to fit it. If you forgot a ResourceQuota scoped to that PriorityClass, one rogue user can preempt the entire cluster. → PriorityClass + PriorityClass-scoped quota together.

32.12 CPU manager state file lost

Node reboots; /var/lib/kubelet/cpu_manager_state is on tmpfs and disappears. Kubelet starts in static mode with no state; existing containers' cpuset.cpus is wrong. Mismatched accounting. → State file lives on disk; verify on every reboot.

32.13 Topology manager rejecting pods on small nodes

single-numa-node policy on a single-socket node = fine. On a 2-socket node where a pod requests 12 cores but each NUMA has only 8 = perpetual rejection. → Use best-effort unless you've verified your workload genuinely benefits from strict locality.

32.14 Underestimating /var/log

Verbose logging + log rotation slack + crash dumps + tmpfile cleanup race + uncaught exceptions printing stack traces = /var/log at 90%. nodefs.available triggers eviction of unrelated pods. → Centralize logs (Loki, Fluent Bit shipping to S3), or split /var/log to its own filesystem.

32.15 Overcommit ratio in autoscaler

Cluster Autoscaler scales nodes based on unscheduled pods (= pending requests). If you've overcommitted (limits >> requests) and CPU usage spikes to limit, the autoscaler doesn't know — it sees requests satisfied. Throttling explodes; latency degrades. The autoscaler should be sized off usage, not requests, for this case. → KEDA + Prometheus-based custom autoscaler for usage-based scaling.

32.16 Init container resource bookkeeping

For QoS computation, every init container counts. For scheduling, the kubelet uses max(max init container request, sum of regular container requests). So an init container with requests.cpu: 4 reserves 4 CPUs at scheduling but releases them after init completes. → Don't over-request on init containers; it inflates apparent reservation.

32.17 Mixed kubeletReserved enforcement

Setting enforceNodeAllocatable: ["pods", "kube-reserved"] puts a hard cgroup memory limit on the kubelet itself. If the kubelet leaks (e.g., a pod with thousands of containers + status manager backlog), the kubelet gets OOM-killed. Node dies. → Default to ["pods"] only unless you have rigorous kubelet memory monitoring.

32.18 hugepages reservation without NUMA awareness

Reserving hugepages at boot via hugepages=N puts them all on NUMA node 0. Pods scheduled to NUMA-1 cores get remote memory access. → Use per-node reservation via /sys/devices/system/node/node*/hugepages/....

32.19 GPU not visible to topology manager pre-1.20

Older device plugins didn't report NUMA topology. Topology manager couldn't align CPU + memory to GPU. → Update device plugins, enable KubeletPodResourcesGetAllocatable and CPUManagerPolicyAlphaOptions=full-pcpus-only for proper alignment.

32.20 ResourceQuota race during admission

Two concurrent pod creates can both pass quota check at admission and both be admitted, briefly exceeding quota. The controller reconciles status.used shortly after. Rare but causes paging if the alert is "quota exceeded > 0". → Allow a small margin (5%) in alerts.

32.21 Pod with only init container counting in QoS

A pod with one init container that has full requests/limits, and no regular containers (rare but legal — restartable init containers / sidecars from 1.28+) — its QoS is computed correctly, but the kubelet's pod cgroup machinery has historically been confused. Pre-1.30, edge cases in cgroup writes. → Use 1.30+ for sidecar-heavy designs.

32.22 oom_score_adj not inheriting

The kubelet sets oom_score_adj on the container's main PID via the CRI. Children inherit at fork time. But: a process that calls prctl(PR_SET_DUMPABLE, 0) or runs setuid loses inheritance. The OOM killer may then pick a child with oom_score_adj=0 instead of the intended -997. → Audit containers that exec setuid binaries; consider hardening with seccomp to forbid those paths.

32.23 In-place resize on a deployment without RollingUpdate

VPA in Auto mode + Deployment with Recreate strategy = VPA evicts all pods at once when resizing. → Use RollingUpdate or VPA Initial mode.

32.24 Memory.high spike during JVM startup

JVMs allocate aggressively during class-loading. With MemoryQoS enabled, hitting memory.high throttles the allocator → JVM startup takes 3× longer. → Disable MemoryQoS for JVM-heavy clusters, or pre-warm with JAVA_OPTS=-XX:+AlwaysPreTouch -Xmx<limit>.

32.25 Eviction thresholds vs node-exporter memory accounting

Node-exporter's node_memory_MemAvailable_bytes includes reclaimable page-cache as "available". The kubelet's memory.available signal can be different (it reads from cgroup root memory.current). Alert on the kubelet's view, not node-exporter's, for eviction-relevance. → kubelet_* metrics, not node_*, for eviction thresholds.

32.26 Cron jobs creating ephemeral pods at quota edge

CronJob with concurrencyPolicy=Allow and quota tight on pods: 100. Two cron runs overlap during a slow execution; one of them can't schedule (pods quota exceeded). Silently fails forever until investigation. → concurrencyPolicy=Forbid or quota with generous headroom.

32.27 kubectl top vs cgroup truth

kubectl top pod uses /metrics/resource (lightweight) which is updated every 10 seconds. Spikes shorter than 10s never appear. For sub-second insight, scrape cAdvisor directly or use kubelet /stats/summary. → Don't make decisions based on kubectl top.

32.28 ServiceAccount token projection counts as ephemeral storage?

No. Projected service account tokens are in tmpfs. But many people think they count, and over-provision ephemeral-storage. → Don't be that person.

32.29 Node drain leaves Guaranteed pods stuck

A drain operation tries to evict pods. Guaranteed pods with PDBs blocking eviction → stuck. → Set PDBs with realistic minAvailable, or --delete-emptydir-data --force (with care).

32.30 Resource fragmentation at the cluster level

20 nodes each with 1 CPU free, but a pod needs 4. Scheduler sees 20 CPUs total free — but can't fit. Pod stays Pending forever. → Bin-pack proactively (scheduler MostAllocated scoring) or use Karpenter to consolidate.


33. TL;DR

Resource management in Kubernetes is two knobs at two layers:

  • requests drive scheduling. They count against the node's Allocatable = Capacity − kubeReserved − systemReserved − evictionHard. requests.cpu also becomes cpu.weight in the cgroup (proportional share under contention). requests.memory becomes memory.min only with the MemoryQoS feature gate.
  • limits drive enforcement by the kernel. limits.cpu → cpu.max (CFS quota/period throttling, default period 100ms). limits.memory → memory.max (hard cap, cgroup OOM on exceed).

The three QoS classes are computed from the spec, not user-declared:

  • Guaranteed: every container has CPU and memory requests == limits. oom_score_adj = -997. Last to die. Eligible for static CPU manager pinning if requests.cpu is an integer.
  • Burstable: anything else with at least one request or limit. oom_score_adj scales 2–999 by memory-request fraction of node capacity.
  • BestEffort: no requests or limits anywhere. oom_score_adj = 1000. First to die. Don't use in production.

The cgroup-v2 tree under /sys/fs/cgroup/kubepods.slice/ has three branches: kubepods-besteffort.slice/, kubepods-burstable.slice/, and direct children for Guaranteed pods. Each pod is a slice; each container is a scope. The kubelet writes cpu.weight, cpu.max, memory.max, optionally memory.min/memory.high, pids.max, and (with managers) cpuset.cpus, cpuset.mems.

CPU throttling is the silent killer. CFS quota throttles a cgroup even if the rest of the node is idle. Multi-threaded apps blow through quota in a fraction of a period and wait. container_cpu_cfs_throttled_periods_total / container_cpu_cfs_periods_total > 25% for 15 minutes is a "fix it now" alert. The fix is usually remove the CPU limit; rely on cpu.weight from requests for fairness.

Memory limits are mandatory in production. Without them, a runaway pod takes down the node. With them, the cgroup OOM contains the blast radius. Set limit ≈ 1.5× p99 working-set.

The static CPU manager pins integer-CPU Guaranteed pods to exclusive CPUs. The memory manager does the NUMA-equivalent. The topology manager is the arbiter that merges hints into a single NUMA placement, with policies none, best-effort, restricted, single-numa-node, at scope container or pod. Strict policies reject pods that can't be NUMA-aligned; use best-effort unless you've measured the benefit.

ResourceQuota caps aggregate resources per namespace. LimitRange sets per-object defaults and bounds. Quotas can be scoped (BestEffort, Terminating, PriorityClass) so that different workload classes have separate budgets. PriorityClass + preemption lets high-priority pods kick out low-priority ones.

Extended resources (nvidia.com/gpu etc.) are integer-only and require requests == limits. Hugepages are reserved at boot and consumed via hugepages-2Mi/hugepages-1Gi resources. Ephemeral storage counts container writable layer + emptyDirs + logs. PID limits (podPidsLimit, default unlimited; pid.available eviction signal) prevent fork bombs.

Eviction is the kubelet's proactive defense against node pressure, distinct from the kernel's reactive OOM. Six signals (memory.available, nodefs.available, nodefs.inodesFree, imagefs.available, imagefs.inodesFree, pid.available), each with hard and soft thresholds. Ranking: BestEffort → Burstable → Guaranteed, then by priority, then by overshoot of request. The eviction manager sets node conditions (MemoryPressure, DiskPressure, PIDPressure) that the scheduler avoids.

In-place pod resize (GA in 1.32) lets you change resources on a running pod via the resize subresource without restarting (unless resizePolicy demands a restart for memory). VPA is the primary consumer.

The right-sizing workflow: measure for a week with generous defaults, set requests.cpu ≈ p95–p99 measured CPU, requests.memory ≈ p99 working-set, limits.memory ≈ 1.5× request, omit limits.cpu for latency-sensitive services. Iterate.

Observability is non-negotiable: alert on OOMKilled, throttling >25%, memory >90% of limit, node MemoryPressure/DiskPressure/PIDPressure, ResourceQuota >90%, eviction count >0. container_cpu_cfs_throttled_periods_total, container_memory_working_set_bytes, container_oom_events_total, kubelet_evictions{eviction_signal=...}, kube_pod_container_resource_* are the five metric families you must scrape.

The sentence to remember: requests buy you scheduling and a CPU share; limits buy you a CPU ceiling and a memory ceiling; CPU ceilings hurt more than they help; memory ceilings are non-negotiable; and the kubelet's eviction manager will save the node before the kernel OOM has to.

Next: chapter 22 covers autoscaling — HPA scales replicas based on the metrics we measured here, VPA changes the requests we set here (via in-place resize), Cluster Autoscaler / Karpenter grow nodes when requests exhaust Allocatable. Every autoscaler is downstream of the choices in this chapter.