KnowSys
Incident24 September 2026· ~14 min

You paid for four CPUs and got throttled at 40%

The dashboard says the pod is using 1.6 of its 4 cores. The kernel says it was throttled in 118 of the last 300 periods. Both are right. They just measure different things.

linuxcgroupskuberneteslatencyscheduling

If you run services on Kubernetes, your pods almost certainly have a CPU limit, set by you or by whoever wrote the Helm chart. It reads like a simple promise: this pod can use up to four cores, and the dashboard tells you how much of that it's using. Under the hood, Linux enforces the limit with the CFS bandwidth controller. It hands the cgroup a budget of CPU time for each period (100 ms by default), and once the budget's spent, every thread in the cgroup waits for the next period, however idle the rest of the machine is.

That's where the dashboard and the kernel stop agreeing. Usage is an average over seconds, and throttling happens inside 100 ms windows, so a pod can look half idle and still spend part of every period frozen. If you've ever raised a CPU limit or added replicas to fix a latency problem that didn't look like a CPU problem, this is probably what you were paying for.

Here's a graph most people who run services on Kubernetes have stared at. A pod with a CPU limit of 4. Usage hovering around 1.6 cores, never touching the line. Each request is about a millisecond of work. And p99 is sitting at 40-something milliseconds, with no deploy, no traffic change and no slow dependency to blame.

I built one of those on purpose while writing this, in a Docker container on my laptop, and the numbers came out like this:

7.0 ms
p99 with no CPU limit
Docker Desktop, linuxkit 6.10, aarch64, 10 vCPUs · 6,061 requests
43.5 ms
p99 with --cpus=4, same work
same machine, same seed, avg usage 1.79 CPUs
118 / 300
Periods throttled
nr_throttled / nr_periods from cpu.stat
45%
Average CPU as a share of the limit
usage_usec over wall time

Six times worse at p99, while the average said there was more than half the limit left over. This post is about why that's not a contradiction, the two separate reasons it happens (one is a design choice and one was a kernel bug), and what the runtimes you actually deploy do to make it worse.

The first three theories

The usual first explanations go roughly in this order, and each one survives longer than it should.

"The node is busy." A noisy neighbour stealing cycles is the natural suspect for tail latency on shared hardware. But a noisy neighbour shows up as the pod getting less CPU than it asked for across the board, and you'd expect p50 to move too. In my run p50 barely changed (1.86 ms to 2.05 ms). Only the tail blew up. That's a hint that something is happening in bursts, not all the time.

"It's the application: GC, a lock, a slow path." Closer. There often is a burst inside the process: a GC cycle, a cache refresh, a fan-out query. But in my program the burst is 800 ms of CPU spread over ten threads: 80 ms of wall time on ten cores. With no limit, that burst costs the handlers almost nothing (p99 7 ms). Put the same burst under a limit and it costs them 36 ms more. So the burst is probably the trigger, but it can't be the delay. Where do the other 36 ms come from?

"We can't be hitting the limit, we're at 45%." This is the one that holds out longest, because the dashboard really does say 45%. Why would you suspect a limit you're nowhere near? It's also, I think, the wrong mental model of what a CPU limit is.

What "4 CPUs" means to the kernel

Chapter 06 covers the basics of the quota, so here's only the part that matters for this story.

A Kubernetes CPU limit of 4 becomes a line in the pod's cgroup:

Shell
$ cat /sys/fs/cgroup/cpu.max
400000 100000

That's 400 ms of CPU time per 100 ms period. It's not four cores. It's a budget of time, and any number of threads can spend it at once. Ten threads on ten cores burn through 400 ms in 40 ms of wall clock. Then every thread in the cgroup is taken off the run queue until the period timer fires and refills the budget, 60 ms later.

Budget is handed out in slices. Each CPU that runs a thread from the group pulls a slice (5 ms by default, sched_cfs_bandwidth_slice_us) from a global pool, and when the pool is empty, that CPU's queue is throttled. Here's the refill, run once per period from an hrtimer (the comments are mine):

kernel/sched/fair.c — __refill_cfs_bandwidth_runtime
torvalds/linux @ v6.10 ↗
C
void __refill_cfs_bandwidth_runtime(struct cfs_bandwidth *cfs_b)
{
	s64 runtime;
 
	if (unlikely(cfs_b->quota == RUNTIME_INF))
		return;
 
	cfs_b->runtime += cfs_b->quota;           /* the period's 400 ms  */
	runtime = cfs_b->runtime_snap - cfs_b->runtime;
	if (runtime > 0) {                        /* we dipped into burst */
		cfs_b->burst_time += runtime;
		cfs_b->nr_burst++;
	}
 
	/* cap: never carry more than quota + burst into a period */
	cfs_b->runtime = min(cfs_b->runtime, cfs_b->quota + cfs_b->burst);
	cfs_b->runtime_snap = cfs_b->runtime;
}

Look at that last min. With cpu.max.burst at its default of 0, unused time doesn't carry over. A period where you idled doesn't buy you anything in the next one. So the quota constrains the average over each 100 ms window, and says nothing at all about how many cores you use inside it.

A request that arrives during those 60 ms waits for the refill, even though its own work is a millisecond and the handler threads were idle. That's most of the extra 36 ms of p99.

Reproducing it

My program is a toy service. Four handler threads take requests from a queue; requests arrive open-loop (Poisson, 200 a second) and each burns 1 ms of CPU. A separate thread starts a parallel phase every 500 ms, spread across par threads, as a stand-in for a GC cycle or a cache rebuild. I sized that phase at ten threads because nproc inside the container says 10, and that's what a runtime that counts CPUs would pick.

Two details matter more than they look. Work is measured on each thread's own CPU clock, so throttling can't make a request do less work. It can only make it wait. And latency is measured from the scheduled arrival time, so a stalled generator can't hide the stall (the coordinated-omission trap).

A 1 ms handler next to a bursty parallel phase
cpp
C++
static double thread_cpu_ms() {
  timespec ts; clock_gettime(CLOCK_THREAD_CPUTIME_ID, &ts);
  return ts.tv_sec * 1e3 + ts.tv_nsec / 1e6;
}
static void burn_cpu(double ms) {          // spin until THIS thread used `ms` of CPU
  double end = thread_cpu_ms() + ms;
  volatile unsigned long x = 0;
  while (thread_cpu_ms() < end) for (int i = 0; i < 1000; ++i) x += i;
}
 
// 4 handlers: pop a scheduled-arrival timestamp, burn 1 ms, record now - arrival
// 1 driver:   every 500 ms, spawn `par` threads that each burn burst_ms / par
ts.emplace_back([&] {
  for (auto next = t0; (next += 500ms) < until;) {
    std::this_thread::sleep_until(next);
    std::vector<std::thread> w;
    for (int i = 0; i < par; ++i) w.emplace_back(burn_cpu, burst_ms / par);
    for (auto& t : w) t.join();
  }
});
 
// generator: Poisson arrivals at 200/s, pushed with their *scheduled* time
for (auto next = t0; next < until;) {
  next += exp_gap(rng);
  std::this_thread::sleep_until(next);
  { std::lock_guard g(m); q.push_back(next); }
  cv.notify_one();
}
Shell
$ docker run --rm --privileged -v "$PWD":/w -w /w throttle-lab ./lab 10 800 500 30
$ docker run --rm --privileged --cpus=4 -v "$PWD":/w -w /w throttle-lab ./lab 10 800 500 30
output

Docker Desktop, linuxkit 6.10.14, aarch64, 10 vCPUs in the VM. Each line is one 30 s run.

Output
cpu.max=max 100000     nproc=10
par=10 burst=800ms/500ms  n=6061  avg CPU 1.79  periods 0    throttled 0    throttled_usec 0
  p50 1.86  p90 3.38   p99 6.99   p99.9 15.53  max 21.20  (ms)
 
cpu.max=400000 100000  nproc=10
par=10 burst=800ms/500ms  n=6061  avg CPU 1.79  periods 300  throttled 118  throttled_usec 27530287
  p50 2.05  p90 18.86  p99 43.54  p99.9 53.15  max 63.36  (ms)

Same work, same arrival sequence (the RNG is seeded), same average usage. The only change is the cpu.max line.

That 118 lines up almost exactly with the arithmetic. There are 60 bursts in 30 seconds, and each one needs 800 ms against a 400 ms budget, so it spans two periods and gets throttled in both. That's 120.

Then there's throttled_usec, which I misread the first time. It says 27.5 seconds of throttling in a 30-second run. That can't be wall time: 118 throttled periods of at most 100 ms each is 11.8 s, tops. It turns out the counter is summed over CPUs. Each per-CPU run queue adds its own frozen time when it's unthrottled:

kernel/sched/fair.c — unthrottle_cfs_rq
torvalds/linux @ v6.10 ↗
C
	raw_spin_lock(&cfs_b->lock);
	if (cfs_rq->throttled_clock) {
		/* one cfs_rq per CPU, all adding into one group-wide total */
		cfs_b->throttled_time += rq_clock(rq) - cfs_rq->throttled_clock;
		cfs_rq->throttled_clock = 0;
	}

So 27.5 s across roughly ten CPUs is about 23 ms per CPU per throttled period. If you've ever graphed container_cpu_cfs_throttled_seconds_total from cAdvisor and seen it exceed wall time, that's probably why. A ratio of container_cpu_cfs_throttled_periods_total to container_cpu_cfs_periods_total is a much easier number to reason about.

The fixes, measured

Partway through writing this, my Linux setup changed to a single shared container, capped at 4 CPUs, that a dozen other jobs were also using. I reran the experiment there with the limit at 2 CPUs (a scaled-down version of the same shape: 300 ms bursts every 500 ms, so the average lands at 0.8 CPUs, or 40% of the limit). Each cgroup was a transient systemd scope with CPUQuota=200%. Other people's load leaks into the tails here, so each row is the median of five 30-second runs.

SetupThrottled periodsthrottled_usecp99 (range of 5)p99.9
No limit, 10 burst threads0 / 00 s11.3 ms (6.0–16.2)21.3 ms
2 CPUs, 10 burst threads59 / 30036.9 s65.8 ms (39.7–83.2)75.2 ms
2 CPUs, 4 burst threads59 / 30016.7 s42.1 ms (5.7–54.0)48.3 ms
2 CPUs, 2 burst threads59 / 3001.2 s6.7 ms (5.9–9.0)13.4 ms
3 CPUs, 10 burst threads59 / 3004.5 s19.2 ms (10.4–50.6)42.9 ms
2 CPUs + cpu.max.burst 200 ms, 10 threads0 / 3000 s6.8 ms (6.6–7.0)13.0 ms

The no-limit row is noisier than the first run was. With nothing capping it, my process plus everyone else's occasionally pushed the whole container past its own 4-CPU cap, and the box-level cpu.stat recorded between 1 and 47 throttled periods per run. Keep that in mind when reading its 11 ms.

Three things in that table I didn't expect.

Burst was the clear winner: zero throttled periods, and the tightest spread of any row, including the unlimited one. With 200 ms of banked quota, a 300 ms burst fits inside a single period, and the average is still enforced over the longer window. Kernel 5.14 or later only, and there's a catch for Kubernetes users I'll come back to.

The other reason: throttled while under quota

Everything above is the quota working as designed. There was also a period, roughly 2018 to 2020, when it didn't, and that's the part most of the "remove your CPU limits" posts were really reacting to.

Ivan Babrou at Cloudflare reported it to LKML in December 2017: services that weren't burning through their quota got throttled anyway, on 4.4, 4.9 and 4.14. He opened kubernetes#67577, "CFS quotas can lead to unnecessary throttling", in August 2018 as a heads-up to Kubernetes.

Then 4.18 made it much worse. Dave Chiluk at Indeed traced the regression to commit 512ac999d275, which fixed clock drift in how per-CPU slices expired. Remember the slices: each CPU pulls 5 ms from the global pool, and when its threads go idle it hands back everything but 1 ms. Before 4.18, a clock-skew bug meant that leftover 1 ms mostly never expired. After the fix, it expired correctly at the end of every period, on every CPU the group had touched.

On one CPU that's noise. On Indeed's 88-core machines it could strand 87 ms of quota per period, in Chiluk's words "87ms or 870 millicores or .87 CPU that could potentially be unusable." The workloads hit hardest were the ones you'd least suspect: highly threaded, mostly idle, waking briefly on lots of different cores. In Indeed's case, Java web services.

His fix removed the expiry entirely: de53fd7aedb1, "sched/fair: Fix low cpu usage with high throttling by removing expiration of cpu-local slices", reviewed by Ben Segall and Phil Auld and merged for 5.4. The commit message is honest about the trade: quota is no longer strictly enforced per period, and a group can overrun by whatever slice is left on each CPU, typically 1 ms or less. It's accurate over longer windows, and highly threaded apps stop getting throttled for time they never used. Indeed reports a worst-case latency on one application going from over two seconds to 30 ms.

It went back to the stable trees. Here's the list worth checking against your fleet:

(Those versions are from Indeed's write-up; I haven't checked each changelog myself.) I can't reproduce this one on the box I have. It's 6.10, and the code that expired slices is gone. If you want to see it, you'd need a 4.18 to 5.3 kernel without the backport, a high core count, and a thread pool that wakes briefly on many CPUs.

Runtimes that count the wrong CPUs

Here's a detail from my own run that I skipped past. Inside the container with --cpus=4, nproc printed 10. A cgroup limit doesn't change what sched_getaffinity returns, so anything that sizes itself from the CPU count sees the whole machine. On a 64-core node with a 4-CPU limit, that's 64 GC threads, 64 worker threads, a 64-way fork-join pool, all sharing 400 ms.

Go is the clearest case. Before 1.25, GOMAXPROCS defaulted to the number of logical CPUs, and the GC's background mark workers target 25% of GOMAXPROCS (gcBackgroundUtilization = 0.25). On that 64-core node a GC cycle wants 16 cores at once. Uber's automaxprocs README has the best single table on this, from one of their services with a 2-CPU quota:

GOMAXPROCSRPSp50p99.9
128,8931.46 ms19.70 ms
2 (equal to quota)44,7150.84 ms26.38 ms
441,0710.57 ms42.94 ms
833,1120.43 ms64.32 ms
24 (the default, host CPUs)22,1910.45 ms76.19 ms

Notice p50 keeps improving as GOMAXPROCS goes up, while throughput halves and p99.9 triples. If you only watch the median you'd tune this in exactly the wrong direction.

Go 1.25 fixed the default. Its runtime now reads the cgroup's cpu.max and uses the smaller of that and the CPU count, and it re-checks periodically, since limits can change under a running pod (release notes, design post). Rounding rules are in the source:

src/runtime/cgroup_linux.go — defaultGOMAXPROCS
golang/go @ go1.25.0 ↗
Go
func defaultGOMAXPROCS(ncpu int32) int32 {
	// GOMAXPROCS is the minimum of:
	//
	// 1. Total number of logical CPUs available from sched_getaffinity.
	//
	// 2. The average CPU cgroup throughput limit (average throughput =
	// quota/period). A limit less than 2 is rounded up to 2, and any
	// fractional component is rounded up.

I checked it on the shared box with a four-line program that prints runtime.NumCPU() and runtime.GOMAXPROCS(0), built with the go1.25.0 toolchain, once with go 1.25.0 in go.mod and once with go 1.24. Each run is its own systemd scope:

Shell
$ systemd-run --scope -p CPUQuota=200% ./gmp125
$ systemd-run --scope -p CPUQuota=200% ./gmp124
Output
shared 4-CPU container, linuxkit 6.10.14 aarch64, go1.25.0 linux/arm64
CPUQuota=200%  nproc=10  gmp125: NumCPU 10 GOMAXPROCS 2
CPUQuota=200%  nproc=10  gmp124: NumCPU 10 GOMAXPROCS 10
CPUQuota=250%  nproc=10  gmp125: NumCPU 10 GOMAXPROCS 3
CPUQuota=250%  nproc=10  gmp124: NumCPU 10 GOMAXPROCS 10
no quota       nproc=10  gmp125: NumCPU 10 GOMAXPROCS 10

That last line surprised me. The whole container is capped at 4 CPUs (its root cgroup's cpu.max reads 400000 100000), yet Go picked 10. It turns out the runtime opens cpu.max for the process's own cgroup and nothing above it (OpenCPU in internal/runtime/cgroup). In a plain docker run --cpus container, or a Kubernetes container, the limit sits on the process's own cgroup and this is fine. If the limit lives on a parent (a systemd slice with CPUQuota=, or this box), Go can't see it.

Two catches. It only kicks in when your go.mod says go 1.25 or later, so upgrading the toolchain alone doesn't change behaviour. And it reads the limit, not the request. Go's reasoning is that a pod with no limit is allowed to use idle CPU beyond its request, so capping it at the request would waste that. That's defensible. It also means that the day you delete the CPU limit on the advice below, GOMAXPROCS silently goes back to the node's core count.

The JVM went through the same thing earlier. JDK-8146115 taught HotSpot to read cgroup quotas in JDK 10 (on by default as -XX:+UseContainerSupport, backported to 8u191), so availableProcessors() returns quota divided by period, rounded up. It also used to treat cpu.shares as a CPU count, so a pod with a small request and no limit saw one CPU. JDK-8281181 stopped that in 18.0.2, 17.0.5 and 11.0.17. If you're on something older, or you pin -XX:ActiveProcessorCount to $(nproc) (Cloud Foundry's buildpack used to), you get the host count and the throttling that comes with it.

Anything else that calls nproc, reads /proc/cpuinfo, or uses std::thread::hardware_concurrency() is in the same boat. How many of the thread pools in your service were sized by a line like that at startup?

"Just remove the limits"

This advice has a real pedigree. Tim Hockin, one of the original Kubernetes engineers, tweeted in 2019: "Always set memory limit == request. Never set CPU limit (for locally adjusted values of 'always' and 'never')." That parenthetical is doing a lot of work, and it usually gets dropped when the tweet is quoted.

Case studies people cite:

Zalando's and Buffer's are from the 2018 to 2020 window, when a lot of fleets were on a kernel that throttled under quota, and Buffer and Omio both point at that kernel bug directly. That's a different problem from the one in my reproduction, and it has a cleaner fix: upgrade.

What removing the limit costs is predictability. Without a limit, a pod's latency depends on how busy its neighbours are, so it's fast in staging and slower in production on the busiest node. Requests still protect you under contention (they become cpu.weight), but they protect a share, not a latency. And as above, Go 1.25 and a modern JVM size themselves from the limit, so deleting it widens every thread pool at the same time.

My own read, and it's only that: for a latency-sensitive service, the default should be no limit, a request sized to real peak usage, and runtime thread counts pinned explicitly. For batch jobs and anything multi-tenant, keep the limit and size the threads to it.

Burst, and why it's not in your pod spec

Linux 5.14 added a third option. Huaixin Chang's patch from Alibaba lets a group bank unused quota up to cpu.max.burst microseconds and spend it later. That's the quota + burst cap in the refill function above. Its kernel documentation describes the trade plainly: it "borrows time now against our future underrun, at the cost of increased interference against the other system users."

It's exactly the right shape for a bursty service that's under quota on average. Kubernetes doesn't expose it. kubernetes#104516, opened in August 2021, is still open. People set it by writing the file from a DaemonSet, through Koordinator's pod annotations, or with Alibaba ACK's CPU Burst policy. Outside Kubernetes, on cgroup v2, you can write it yourself:

Shell
# the cgroup of a systemd-managed service
echo 200000 > /sys/fs/cgroup/system.slice/myapp.service/cpu.max.burst
cat /sys/fs/cgroup/system.slice/myapp.service/cpu.stat   # nr_bursts, burst_usec

What I can't explain

Look at the spread in the 2-CPU, ten-thread row. Five runs, same binary, same seed, and each one throttled in exactly 59 or 60 periods. But p99 ranged from 39.7 ms to 83.2 ms, and throttled_usec from 22.2 s to 46.3 s, moving together. The first --cpus=4 run has the same smell: it predicts about 600 ms of per-CPU throttling for each throttled period (ten CPUs, 60 ms each), and it measured 233.

My best guess is phase. The burst fires every 500 ms from the start of the program, the period timer fires every 100 ms from whenever the group first ran, and since 500 is a multiple of 100 the offset between them is fixed for a whole run and random between runs. A burst that starts just after a refill gets throttled for most of the period; one that starts late spills into the next refill and barely stalls. That would explain both numbers. I haven't proven it. The experiment that would settle it is jittering the burst start inside each run, or logging throttle_cfs_rq and unthrottle_cfs_rq with bpftrace next to the burst timestamps. If it's right, it means the same service on two pods can have very different tails from the same limit, for reasons nobody will ever find in a dashboard.

What to do on Monday

The machinery behind this story
More field notes