KubeGlanceDownload
Kubernetes requests vs usage: finding reserved but idle CPU and memory

operations

Kubernetes requests vs usage: finding reserved but idle CPU and memory

Compare what pods reserve with what they use over a week, in PromQL: the rate() trap, p95 vs peak, safe requests, and which pods get evicted first.

· 11 min read

A cluster can be full and idle at the same time. The scheduler has no free room left for your next pod, and kubectl top shows the nodes loafing at a tenth of their CPU. Both are true, because they measure different things: the scheduler counts requests, the graphs show usage, and nothing in Kubernetes compares the two for you.

This post does that comparison with PromQL. You get the room a namespace reserves against what it used over a week, a request you can put in a manifest, and the one case where the gap goes the other way and a pod ends up first in line for eviction. It also covers the query that looks right and is wrong, which we shipped ourselves before a blog post caught it.

What each number means

Requests decide where a pod lands. Usage decides nothing until memory runs short. Then usage above the request decides who leaves first.

A request is a reservation. The scheduler subtracts it from a node’s allocatable room whether or not the pod ever uses it. We covered how the CPU half of that turns into cgroup weights in CPU requests vs limits. So every node’s room splits three ways, not two:

  • In use: reserved and used.
  • Reserved, idle: reserved and not used. The scheduler treats this as full. Trimming requests gives it back.
  • Unreserved: never asked for. The scheduler can already place pods here.

A single “utilisation” percentage blurs the middle band with the last one, and those two call for opposite fixes. Idle reservation means lowering requests. Unreserved room means you have more nodes than your requests need. That is why a drain can strand pods on a cluster that looks idle.

Step 1: what each namespace reserves

kube-state-metrics exports every container’s requests. Filter to running pods; otherwise completed Jobs keep contributing requests they no longer hold:

sum by (namespace) (
  kube_pod_container_resource_requests{resource="cpu"}
  * on (namespace, pod) group_left ()
  (kube_pod_status_phase{phase="Running"} == 1)
)
kube-system   1.05
monitoring    0.12

That output is from a test cluster running the stock kube-prometheus-stack, and it shows the first surprise: most of that chart sets no requests at all. A pod with no request uses room the scheduler thinks is empty. It does not show up as waste here, and it shows up as risk further down.

Step 2: usage over a week, and the query that lies

The obvious query for “average CPU per pod over seven days” is a rate over seven days:

# Wrong whenever there is less than 7 days of data
sum by (namespace, pod) (
  rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[7d])
)

It returns plausible numbers, and they can be off by any factor. Prometheus divides a rate by the whole range you asked for, not by the part that holds samples. Take a pod that started yesterday: its seven-day “average” is its real average divided by seven. On a backend with two days of retention, every pod reads 3.5 times low. Our test cluster had 1.7 hours of data, which put this query about a hundred times low:

                              rate(...[7d])   mean of 5m rates
kube-apiserver                0.00041         0.0398
etcd                          0.00017         0.0162
prometheus                    0.00012         0.0116

We had this exact query in the app’s Efficiency screen. On a production cluster with two days of retention, it reported 2% of reserved CPU in use; the true figure was 8%. The low number looked believable, which is why it survived. It turned up while we were writing this post, when a seven-day average came out a hundred times smaller than its own p95.

VictoriaMetrics and Thanos behave the same way; we checked both. This is not a quirk of one backend.

The fix is to average short rates over the steps where the pod actually has data:

avg_over_time(
  sum by (namespace, pod) (
    rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
  )[7d:5m]
)

It is a subquery: every 5 minutes for 7 days, take a 5-minute rate, then average the rates that exist. A pod that has lived for one day is averaged over that day. Memory needs none of this, because avg_over_time on a gauge already averages only the samples it has.

Then ask how much history there actually is, because “7 days” in the query is a request, not a promise:

time() - min_over_time(
  timestamp(count(container_memory_working_set_bytes{container!="", container!="POD"}))[7d:5m]
)
6181.29

That is seconds: 1.7 hours of a requested week. Put that next to any average you report. A number without its coverage invites the wrong conclusion.

With both halves, namespace efficiency is a division:

sum by (namespace) (
  avg_over_time(sum by (namespace, pod) (
    rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
  )[7d:5m])
)
/
sum by (namespace) (
  kube_pod_container_resource_requests{resource="cpu"}
  * on (namespace, pod) group_left () (kube_pod_status_phase{phase="Running"} == 1)
)
monitoring    0.249
kube-system   0.074

Step 3: a request you can write down

An average tells you there is slack. It does not tell you the request to set, because a workload that averages 0.1 cores and spends lunchtime at 0.8 needs the 0.8. The two resources fail differently, so they get different high-water marks:

  • CPU: the 95th percentile of 5-minute rates. CPU above the request is not fatal. The pod competes for spare cycles and, at worst, slows down. Covering the p95 lets the rare spike borrow instead of reserving for it all week.

    quantile_over_time(0.95,
      sum by (namespace, pod) (
        rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
      )[7d:5m]
    )
    
  • Memory: the peak. Memory is not compressible. Crossing the limit is an OOMKill, and crossing the request under pressure is a place near the front of the eviction queue (next section).

    sum by (namespace, pod) (
      max_over_time(container_memory_working_set_bytes{container!="", container!="POD"}[7d])
    )
    

Requests are per pod, and replicas are not equal. A leader, a shard with the hot keys or the pod that happens to take the cron traffic will run hotter than the rest. So size from the busiest pod’s figure, not the workload’s mean. Add headroom (we use 15%) and round to something a person would type: 170m, not 0.16837.

Two rules keep this from being confidently wrong:

  1. No suggestion without a full day of data, counted from the youngest pod. A rollout restarts every pod’s clock. Two hours of history after a deploy misses the nightly batch, and a request sized from it will be too small.
  2. Sample at the same step you average over. The p95 and the average above share one [7d:5m] subquery, so the p95 can never sit below the average because of different sampling. If it does, the distribution really is spike-shaped: mostly idle, with a few bursts.

Step 4: memory above its request, the part that is not about saving

The gap can also run the other way. A pod whose memory sits above its request is not saving you anything. It is using room nobody reserved for it, and when that room runs out, it is the first to go. The kubelet’s eviction order, from the node-pressure eviction docs:

  1. Whether the pod’s resource usage exceeds requests
  2. Pod Priority
  3. The pod’s resource usage relative to requests

And, less well known: “The kubelet does not use the pod’s QoS class to determine the eviction order.” A Burstable pod that stayed within its request outlives one that exceeded it, whatever the QoS class.

sum by (namespace, pod) (container_memory_working_set_bytes{container!="", container!="POD"})
> on (namespace, pod)
sum by (namespace, pod) (kube_pod_container_resource_requests{resource="memory"})
monitoring  prometheus-kps-kube-prometheus-stack-prometheus-0   355606528

Prometheus itself, at 339 MiB against a 256 MiB request, is the pod this node evicts first. Meanwhile its sidecars request nothing. This comparison can’t see pods that request no memory at all. The query finds no series to match them against, so check those separately: to the scheduler they weigh nothing, yet they hold real memory.

On the production cluster, this list was more useful than the savings list. One rendering StatefulSet requested 1 GiB per pod and peaked near 4. It is exactly the pod you want to keep, and it would have been the first one evicted.

What it looked like on a real cluster

The production cluster from earlier, over the two days its Prometheus retains:

ReservedUsed (average)In useReserved, idle
CPU5.09 cores0.39 cores8%4.70 cores
Memory18.4 GiB10.3 GiB56%8.2 GiB

The pattern is typical. CPU requests are guesses, made once, and they are high, because a pod that gets too little CPU slows down and someone notices. Memory runs closer to the line, because a request that is too low shows up as an eviction. So the savings are in CPU, while the risk sits in a few memory-heavy workloads. Split that way, the work is two short lists instead of a cluster-wide right-sizing project.

Doing it without writing the PromQL

KubeGlance 1.2 has an Efficiency screen that runs these queries for you. It finds the cluster’s Prometheus, VictoriaMetrics or Thanos on its own, through the API server’s service proxy (how that works is in querying Prometheus through the Kubernetes API), and it states how much history the backend holds instead of pretending it has a week:

KubeGlance Efficiency on iPhone: a week's average per cluster split into in use, reserved but idle and unreserved, with coverage stated and memory-above-request risks listed first
Each cluster's room split three ways, with memory above its request listed before any savings.

On a larger screen, namespaces are ranked by idle reservation. Workloads get a suggested per-pod request only where a full day of history backs it. A bursty workload is told to raise its request rather than being flagged as wasteful:

KubeGlance Efficiency on iPad: namespaces ranked by idle reservation and workloads with average bars, p95 marks and suggested requests such as 300m to 170m per pod
Averages as bars, p95 as a mark, and the request that would cover the busiest pod.

The short version

  1. Requests come from kube-state-metrics, restricted to running pods. Usage comes from cAdvisor.
  2. Never rate(x[7d]) for an average. Average 5-minute rates with avg_over_time(...[7d:5m]), and report the coverage next to it.
  3. CPU requests from the busiest pod’s p95, memory from its peak, plus headroom, and only after a full day of data.
  4. Check memory above request before you cut anything. Those pods are evicted first, QoS class notwithstanding.

See reserved against used for every cluster, with suggested requests and eviction risks, found without configuration. KubeGlance is a native Kubernetes client for iPhone and iPad, with a full Mac app on the same core.

Get KubeGlance
#requests #prometheus #promql #eviction #capacity #cost

The Kubernetes dashboard that fits in your pocket

KubeGlance is a native Kubernetes client for iPhone and iPad — the real dashboard, not a companion — with a full Mac app on the same core. Pods, workloads, logs and events, straight from your kubeconfig.

Download KubeGlance