KubeGlanceDownload
CPU requests vs limits: what each one writes to the cgroup

concepts

CPU requests vs limits: what each one writes to the cgroup

Requests become cpu.weight, limits become cpu.max. One is a share of a contended CPU, the other is a hard ceiling enforced whether or not the node is busy.

· 12 min read

A CPU request and a CPU limit are not two settings for the same thing. They are written to two different files, read by two different parts of the kernel’s scheduler, and they take effect at different times.

  • requests.cpu becomes cpu.weight. It is a relative share, and it only does anything when the CPU is contended.
  • limits.cpu becomes cpu.max. It is a bandwidth cap, and it is enforced in every scheduling period whether the node is busy or completely idle.

That second sentence is the one people get wrong, and it explains the symptom that starts most of these arguments: a pod being throttled on a node that is almost entirely idle.

The fastest check

The values are on the node, not in the API. On a cgroup v2 system:

kubectl get pod <pod> -o jsonpath='{.metadata.uid}{"\n"}'

Then on that node, under the pod’s slice:

cat /sys/fs/cgroup/kubelet.slice/kubelet-kubepods.slice/*/kubelet-kubepods-*pod<uid>.slice/cpu.max
50000 100000

Two numbers: quota and period, in microseconds. 50000 out of every 100000 is half a CPU, which is what limits.cpu: 500m means. max 100000 in the first field means no limit is set.

What the kubelet actually writes

Measured on Kubernetes 1.36.1, cgroup v2, containerd. Six pods, requests only, no limits, reading cpu.weight from each pod’s own cgroup:

requests.cpuCPU sharespod cpu.weight
50m512
100m1024
250m25610
500m51220
1000m102439
2000m204879

The conversion is the documented cgroup v1 to v2 mapping, and it reproduces every row exactly:

shares = millicpu * 1024 / 1000
weight = 1 + ((shares - 2) * 9999) / 262142

The important property is not the formula, it is that the result is proportional to your request. Double the request, double the weight. That is the whole mechanism: when two pods want the CPU at the same instant, the kernel splits it in the ratio of their weights. A pod requesting 500m next to a pod requesting 250m gets two thirds of a contended core, not “500 millicores”.

Requesting more than you use is therefore not free and not harmless — it is a claim on a share you will win during every contention event, at your neighbour’s expense, and it is a claim the scheduler reserves node capacity for.

One detail worth knowing if you go looking: the weight on the container cgroup inside the pod is derived by the container runtime with a different, non-linear conversion — the 250m container above reads 35, not 10. That number only governs how the pod’s own share is divided between its containers. The pod-level weight is the one that competes with other pods, and it is the linear one.

Weight decides how a contended level is split between its siblings. Every level splits independently.

Why you get throttled on an idle node

cpu.max is a quota per period, and the period is 100 milliseconds by default. A container with limits.cpu: 500m may use 50ms of CPU time in each 100ms window. When it has spent that, the kernel stops scheduling it until the window rolls over — regardless of how much CPU the node has spare.

For a service handling a request in a burst of work, that is the failure. Say a request takes 30ms of CPU across four threads. That is 120ms of CPU time, which exceeds a 50ms quota in the first window; the work finishes two windows later. Latency goes from 30ms to 230ms with most of the node idle, and the CPU usage graph shows nothing wrong, because the container really did only use 500m on average.

This is why throttling is invisible to utilisation dashboards. Average usage is not the quantity being enforced. The quantity being enforced is usage per 100ms.

Requests reach the scheduler and the kernel's weight; limits reach only the kernel's bandwidth controller.

The argument, resolved

Two claims circulate, and both are half-right:

“Requests protect you from noisy neighbours.” True under contention, false otherwise. Weight only applies when more runnable work exists than CPU. On an uncontended node, a pod with a 100m request can use every core on the box. If your concern is a neighbour stealing CPU during a spike, requests are exactly the right tool. If your concern is your own service using more than you budgeted, requests do nothing.

“Only limits give you a real guarantee.” False, and backwards. A limit guarantees a ceiling, not a floor. It cannot give you CPU; it can only take it away. The floor comes from the request, via the weight and via the scheduler refusing to place more requested CPU on a node than it has.

The reason the “stop using CPU limits” position keeps winning arguments is that for a latency-sensitive service, the ceiling costs you tail latency and buys you nothing you cannot get from requests. The reason it is not universal advice is that it assumes everything on the node has honest requests. It does not hold on a shared cluster where one team’s runaway process can saturate a node and every other workload discovers what “compressible resource” means during an incident.

A defensible default: set requests always, set CPU limits only where you need predictability more than latency — batch work, untrusted tenants, anything you are billing for. Memory is the opposite case and always needs a limit, because memory cannot be throttled, only reclaimed by killing something. That asymmetry is why OOMKilled and exit code 137 exists as a category and “CPUKilled” does not.

Checking whether you are actually throttled

Not from kubectl. The metrics API carries CPU and memory usage only — there is no throttling field in metrics.k8s.io, which is why kubectl top disagrees with your dashboard. The counter you want is from cAdvisor, via Prometheus:

rate(container_cpu_cfs_throttled_periods_total[5m])
/
rate(container_cpu_cfs_periods_total[5m])

That ratio is the fraction of 100ms windows in which the container hit its quota. Look at it per container, not per pod, and treat a sustained value above a few percent on a latency-sensitive service as a real finding. The raw throttled_seconds_total counter is much less useful on its own, because a large number of very short throttles and a small number of long ones read the same.

The runtimes have opinions about this

Language runtimes read the cgroup, and what they read is the limit, not the request. Versions matter here and the defaults have moved recently.

Go 1.25 and later default GOMAXPROCS from the cgroup CPU bandwidth limit when that is lower than the machine’s logical CPU count, and update it if the limit changes. The release notes are explicit that the runtime does not consider CPU requests. So a Go service with no CPU limit still sizes its scheduler from the node’s core count — 128 threads on a 128-core node inside a container entitled to one. Set the limit, or set GOMAXPROCS yourself. The behaviour can be turned off with GODEBUG=containermaxprocs=0.

Do not assume the memory side followed. GOMEMLIMIT is still manual: the proposal to default it from the cgroup memory limit, golang/go#75164, is open and unshipped as of Go 1.25. Readers who know the GOMAXPROCS headline routinely assume both changed. Only one did.

The JVM has derived availableProcessors from the cgroup CPU limit since container support landed in JDK 10 (JDK-8146115), and GC thread counts and the common ForkJoin pool are sized from it. A limit of 1 gives you a serial collector and a one-thread pool, which is occasionally the surprise behind “the same JAR is slower in Kubernetes”.

Node.js does not size libuv’s threadpool from the cgroup at all — the CLI documentation states plainly that the pool has a fixed size, changed only by UV_THREADPOOL_SIZE. It therefore does not shrink when you lower the limit, and the quota shows up as event-loop lag rather than as a smaller pool.

Preventing the surprises

  • Set requests from observed usage, not from the limit. The request is what the scheduler packs with and what wins contention. A request of 100m on a service that steadily uses 800m is a promise you break on every busy node.
  • If you set a CPU limit, do not set it near the request. The gap is the burst headroom, and a limit equal to the request means you are throttled the moment you need more than average.
  • Do not copy limits between environments. The same limit on a 4-core dev node and a 96-core production node produces very different runtime behaviour, because the runtimes derive thread counts from the limit and the fallbacks from the node.
  • Check the QoS class you ended up in. Requests equal to limits on every container is Guaranteed; anything else is not, and the difference decides who gets evicted first.
kubectl get pod <pod> -o jsonpath='{.status.qosClass}{"\n"}'

For what happens when the scheduler cannot find room for the requests you set, see 0/3 nodes are available.

#cpu #requests #limits #cgroups #throttling

The Kubernetes dashboard that fits in your pocket

KubeGlance is a native Kubernetes client for iPhone and iPad — the real dashboard, not a companion — with a full Mac app on the same core. Pods, workloads, logs and events, straight from your kubeconfig.

Download KubeGlance