KnowSys
Field guide23 September 2026· ~16 min

The dashboard said 60%. The OOM killer disagreed.

Your container has at least five different memory numbers, they disagree by hundreds of megabytes, and the kernel, the kubelet and your Grafana panel each trust a different one.

linuxmemorycgroupskubernetes

Every container you run with a memory limit, in Kubernetes or plain Docker, is one number away from being killed. When a cgroup reaches its limit and the kernel can't reclaim enough to stay under it, the kernel picks a process and sends it SIGKILL. Kubernetes reports that as OOMKilled, exit code 137. There's no exception to catch and no last log line from your app. It's just gone.

So the number you watch matters. Most teams watch a Grafana panel fed by cAdvisor and assume it's the number the kernel checks. It usually isn't, and when the two disagree, the kernel's the one that acts. Knowing which number is which is the difference between sizing a limit properly and bumping it every time a pod dies.

In November 2022 someone put a 500 MiB limit on a single-container deployment, and it started getting OOMKilled. Prometheus and K9s both said the container sat around 300 MiB and never got near the limit. So they SSH'd into the node, found the pod's cgroup, and read the kernel's own counter: memory.usage_in_bytes said 501,346,304, and max_usage_in_bytes said 524,292,096. Its limit was 524,288,000. That's kubernetes/kubernetes#114142, and it was closed two and a half years later with a suggestion to ask AWS support.

A dashboard at 60% and a kernel at 100% aren't a contradiction. They're two different questions, and most of us have only ever asked one of them. I spent a couple of days reproducing each way the numbers split apart, on Docker Desktop's Linux VM (linuxkit 6.10.14, aarch64), with docker run --memory=256m for the early runs and a child cgroup with memory.max set by hand for the later ones. One of the reproductions I still can't fully explain. It's near the end.

First theory: it's a leak, and the scrape missed the spike

This is what everyone reaches for, and sometimes it's right. Prometheus scrapes every 15 or 30 seconds; an allocation burst that lasts two seconds won't show up in any graph, and the kernel kills in microseconds. If your process holds a big transient buffer while parsing a request, you can die between two points on a graph that looks flat.

But look at what #114142 actually shows. That cgroup counter wasn't spiking. It sat at the limit, steadily, while the dashboard sat at 300. A spike doesn't explain a gap that stays open for hours, so whatever the dashboard is measuring, it's missing something that's there all the time.

Second theory: RSS is the truth, and the dashboard is wrong

Next, people kubectl exec in and run ps or top, and both report RSS. If RSS agrees with the dashboard, people conclude the kernel is confused. It isn't. RSS is the wrong ruler for a cgroup, and it's wrong in both directions.

It over-counts shared memory. I mapped the same 200 MB tmpfs file into two processes and read /proc/PID/smaps_rollup for each:

Output
pid 12: Rss: 205612 kB  Pss: 102600 kB  Pss_Shmem: 102400 kB
pid 14: Rss: 205608 kB  Pss: 102640 kB  Pss_Shmem: 102400 kB
cgroup memory.current = 239M   (shmem=200M)

Sum the RSS column and you get 401 MB for 200 MB of actual pages. PSS (proportional set size) splits each shared page among the processes that map it, so the two PSS values sum to about 200 MB. That's right. Postgres is where you meet this: every backend maps shared_buffers, and a per-process RSS view of a Postgres pod can add up to several times the machine.

It also under-counts, badly. RSS only sees pages mapped into a process's page tables. Page cache from read(), files in a tmpfs nobody has mapped, and kernel memory are all invisible to it. Your cgroup sees all of them, and the next few sections are each one of those.

What the cgroup is actually adding up

Linux's cgroup v2 documentation lists what memory.current charges, and it's broader than most people assume:

One rule in that document explains a lot of weirdness on shared nodes: "A memory area is charged to the cgroup which instantiated it and stays charged to the cgroup until the area is released." First touch wins. If your sidecar reads a 2 GB log file first, it owns those cache pages, and your main container gets them for free.

Page cache: charged, and usually harmless

My first reproduction is the one that's supposed to be scary. I wrote a 1 GB file outside the limit, dropped caches, and cat'd it to /dev/null from a container limited to 256 MB. working_set here is the Kubernetes definition, memory.current - inactive_file (more on that in a minute).

Output
start                  current=   2M working_set=   1M | file=2M   inactive_file=1M
1 GB file, read once   current= 255M working_set=   4M | file=253M inactive_file=251M
1 GB file, read again  current= 255M working_set=   5M | file=251M inactive_file=250M
memory.events: low 0 high 0 max 7196 oom 0 oom_kill 0

Nothing died. memory.current hit the limit 7,196 times and every time the kernel evicted clean cache pages and carried on. This is the healthy case, and it's why "usage at 100%" by itself tells you almost nothing about a file-reading workload. Chapter 08 has the page cache itself, and why it's worth 87× on reads; here it's just a tenant that pays rent on demand.

That's the metric that lies high. Working set lies in a direction I didn't expect.

container_memory_working_set_bytes, and the LRU underneath it

Kubelet doesn't evict on memory.current, because that would evict every pod that ever read a file. It uses the working set, and cAdvisor computes it like this (container/libcontainer/handler.go at v0.49.1):

Go
inactiveFileKeyName := "total_inactive_file"
if cgroups.IsCgroup2UnifiedMode() {
	inactiveFileKeyName = "inactive_file"
}
 
workingSet := ret.Memory.Usage
if v, ok := s.MemoryStats.Stats[inactiveFileKeyName]; ok {
	if workingSet < v {
		workingSet = 0
	} else {
		workingSet -= v
	}
}
ret.Memory.WorkingSet = workingSet

Usage minus inactive file pages. Everything else counts, including active_file: page cache the kernel has seen used more than once and will reclaim second. The node-pressure eviction docs say kubelet excludes inactive_file "as it assumes that memory is reclaimable under pressure." The implication is that active page cache is treated as if it weren't. kubernetes/kubernetes#43916, "kubelet counts active page cache against memory.available (maybe it shouldn't?)", was opened in March 2017 and is still open.

So I read a 150 MB file three times. That should promote it to the active list and inflate the working set, probably. It didn't:

Output
lru_gen=0x0001
read() x3              current= 153M working_set=   1M | inactive_file=151M active_file=0M
mmap touch x3          current= 153M working_set= 152M | inactive_file=1M   active_file=150M
lru_gen=0x0000
read() x3              current= 153M working_set= 152M | inactive_file=0M   active_file=151M
mmap touch x3          current= 153M working_set=   2M | inactive_file=150M active_file=1M

This kernel runs with MGLRU on (/sys/kernel/mm/lru_gen/enabled is 0x0001). Under MGLRU, repeated read() calls left the file "inactive" and the working set at 1 MB, while touching the same file through mmap made all of it "active". Then I switched MGLRU off and ran it again, and the two access patterns swapped. Same file, same kernel, same 153 MB in memory.current, and a working set of either 1 MB or 152 MB depending on the LRU implementation and on whether your process uses read() or mmap. 152 of 256 is 59%.

I don't have a full account of why each implementation classifies these the way it does; roughly, the classic LRU promotes on a second read() and only notices mapped accesses during a reclaim scan, and MGLRU does close to the opposite. What matters on Monday is that working set is a heuristic about an LRU, and the LRU changed. If your nodes moved to a kernel with MGLRU on by default (a build option, CONFIG_LRU_GEN_ENABLED), your eviction behaviour for mmap-heavy workloads like Prometheus's own TSDB or anything using LMDB may have moved with it.

Dirty pages: the case where cache does kill you

Here's the reproduction that surprised me most. Before the read test, I created the 1 GB file from inside the 256 MB container with dd if=/dev/zero of=/big bs=1M count=1024. It died with exit 137. Here's the kernel's report from dmesg, trimmed:

Output
dd invoked oom-killer: gfp_mask=0x101cca(GFP_HIGHUSER_MOVABLE|__GFP_WRITE), order=0
  ...
  ext4_da_write_begin+0xac/0x298
memory: usage 262144kB, limit 262144kB, failcnt 2357
Memory cgroup stats for /docker/1d9d1984b96e...:
anon 1638400
file 257724416
file_dirty 257642496
file_writeback 0
[  pid  ]   uid  tgid total_vm      rss rss_anon rss_file ... name
[ 118400]     0 118400      975      586       32      554 ... bash
[ 118417]     0 118417      843      525      256      269 ... dd
Memory cgroup out of memory: Killed process 118400 (bash) total-vm:3900kB, anon-rss:128kB ...

Read it slowly. Anonymous memory is 1.6 MB. File pages are 257.7 MB and 257.6 MB of those are dirty, with zero under writeback. Clean page cache can be dropped instantly; dirty pages have to reach disk first, and nothing had started writing them. The OOM killer then picked the process with the largest footprint, which was bash at 586 pages, not dd at 525. bash was PID 1, so the whole container went. It reproduced on overlayfs and on a plain named volume.

This isn't new. Kernel bugzilla 207273, filed in April 2020, is titled "cgroup with 1.5GB limit and 100MB rss usage OOM-kills processes due to page cache usage after upgrading to kernel 5.4." The reporter's database containers were dying while restoring backups with xtrabackup, and "all the memory is consumed by file_dirty." Michal Hocko's reply suggested cgroup v2, since it "has much better memcg aware dirty throttling implementation so such a large amount of dirty pages doesn't accumulate in the first place."

If you run database restores, log shippers or anything that streams large files to disk inside a tight limit, this is your OOM. RSS is flat, and so is anon.

/dev/shm: memory that outlives its owner

tmpfs pages are charged like page cache and can't be dropped, because there's no file behind them to reload from. Without swap they're as permanent as heap. I wrote to /dev/shm (with --shm-size=1g, so tmpfs itself wouldn't fill) in 1 MB chunks, printing my own RSS alongside the cgroup's shmem counter:

Output
wrote   64 MB  my RSS  1928 kB  cgroup shmem   64 MB
wrote  128 MB  my RSS  2056 kB  cgroup shmem  128 MB
wrote  192 MB  my RSS  2056 kB  cgroup shmem  192 MB
Killed
-rw------- 1 root root 264892416 Sep 25 13:12 /dev/shm/blob
after the kill   current= 255M working_set= 255M | anon=0M file=253M shmem=252M inactive_file=0M
memory.events:   max 168 oom 1 oom_kill 1
dmesg: Memory cgroup out of memory: Killed process 141129 (shmw) anon-rss:1024kB, file-rss:1032kB

That process had 2 MB of RSS. After it died, the cgroup was still at 255 MB, because the file was still there, and killing processes can't delete a file. Anything else that allocated in that container would've been the next victim. Notice also that working_set equals current: shmem lives on the anon LRU, so inactive_file doesn't subtract any of it.

In my first attempt at this I'd set the shell's oom_score_adj to -1000 and forgot that children inherit it. With nothing killable, the kernel logged nine OOM events and zero kills, and write() simply came back short. That's a different failure mode worth knowing about: an unkillable cgroup at its limit turns allocations into errors.

Where you meet this for real:

Kernel memory nobody's process owns

Last category: the kernel's own objects. I ran a program that called stat() on a million paths that don't exist, inside a child cgroup:

Output
before              current=  0M | slab_reclaimable=0M
1M failed stat()s   current= 18M | slab_reclaimable=17M
kernel 18812928   slab 18773376

Each failed lookup leaves a negative dentry, the kernel's cached "no such file" answer. After the program exited there were no processes in the cgroup at all, and it still held 18 MB. It's reclaimable, but working_set counts it anyway since it isn't inactive_file. Anything that probes for files along a search path, like an interpreter resolving imports, can grow this steadily. TCP socket buffers are charged the same way and show up as sock in memory.stat, so a proxy holding lots of slow connections carries some too.

memory.high, memory.max, and three different killers

Everything so far has been about memory.max, the hard limit. Reach it, fail to reclaim, and the cgroup OOM killer runs:

Output
 208 MB touched  t=0.05s
 224 MB touched  t=0.07s
 240 MB touched  t=0.07s
exit 137
memory.events: max 40 oom 1 oom_kill 1
dmesg: oom-kill:constraint=CONSTRAINT_MEMCG ... task=anon

constraint=CONSTRAINT_MEMCG is how you tell this apart from the global OOM killer: that one shows CONSTRAINT_NONE and runs when the whole machine is out of memory. They don't choose from the same pool. A cgroup killer only looks inside the cgroup that hit its limit, and a global one looks at every process on the node, and a pod well under its own limit can still die when a neighbour without one exhausts the node.

memory.high is the other knob. Kernel docs say going over it means processes "are throttled and put under heavy reclaim pressure," and that it "never invokes the OOM killer." I set memory.high=64M and memory.max=128M and allocated 96 MB of anonymous memory, with no swap:

Output
  48 MB touched  t=0.00s
  64 MB touched  t=0.00s
  80 MB touched  t=27.30s
exit 124            (killed by my 60 s timeout, not by the kernel)
memory.events:   high 537 max 0 oom 0 oom_kill 0
memory.pressure: some avg10=91.76 avg60=53.88 ...  full avg10=91.76 avg60=53.88 ...

Those first 64 MB took no measurable time. Another 16 MB took 27.3 seconds, and 28.1 seconds on a repeat; a third run didn't get past 64 MB within 40 seconds (these are on a shared 4-CPU container, so treat the spread as noise). Anonymous memory can't be reclaimed without swap, so all the throttling can do is make the process wait. It's probably the right tool when there's something to reclaim, and a very slow hang when there isn't. Kubernetes' Memory QoS feature, alpha since 1.22 and reworked in 1.27, sets memory.high from your limits, so this isn't hypothetical if you've turned it on.

Look at that pressure line. PSI (pressure stall information, documented here) reports the share of time tasks were stalled on memory. some means at least one task was waiting; full means all of them were. 91.76% over ten seconds is a process doing almost nothing but waiting for memory. My page cache test, by contrast, read 1 GB through a full cgroup and PSI stayed at 0.00. That's the difference between a cgroup that's full and a cgroup that's in trouble, and it's the number systemd-oomd and Meta's oomd act on.

Then there's a third killer, and it isn't the kernel at all. Kubelet evicts pods when node-level memory.available, computed from the working set, crosses its threshold. It works in userspace, and it's why a pod can be evicted with the reason "The node was low on resource: memory" while the kernel never logged an OOM. And since Kubernetes 1.28 set memory.oom.group on cgroup v2, a kernel OOM kills every process in the container, not just the biggest. Preferred Networks wrote up how that broke their ML platform and how they got singleProcessOOMKill into kubelet in 1.32 to opt back out.

Allocators hold on to what you freed

Everything above is the kernel counting correctly. This part is user space keeping memory the program thinks it gave back (chapter 05 covers it in general). Two runtimes deserve a specific look.

glibc gives each thread its own arena under contention, up to a limit it computes from the CPU count (malloc/arena.c, glibc 2.39):

C
if (mp_.arena_max != 0)
  narenas_limit = mp_.arena_max;
else if (narenas > mp_.arena_test)
  {
    int n = __get_nprocs ();
 
    if (n >= 1)
      narenas_limit = NARENAS_FROM_NCORES (n);

NARENAS_FROM_NCORES(n) is n * 8 on 64-bit. __get_nprocs() counts CPUs the kernel reports as online, and a CFS quota doesn't change that. My shared lab container I used has a cpu.max of 400000 100000 (4 CPUs), and nproc still says 10. glibc will let a 4-CPU pod grow 80 arenas.

Each arena returns memory to the kernel only from its top. So I wrote a program where eight threads take turns: each allocates 32 MB in 4 KB chunks, frees all of them except the last one, and hands over. At the end, 32 KB is allocated:

Output
default:            still allocated: 32 KB   RSS  258 MB
MALLOC_ARENA_MAX=2: still allocated: 32 KB   RSS   66 MB
MALLOC_ARENA_MAX=1: still allocated: 32 KB   RSS   33 MB
default:            still allocated: 32 KB   RSS  258 MB   after malloc_trim(0): 2 MB

Identical three times (glibc 2.39, same shared container). One pinned 4 KB chunk per arena holds its 32 MB hostage, eight times over. Heroku saw this across Ruby apps and, as of September 24th, 2019, sets MALLOC_ARENA_MAX=2 by default on new apps. Nate Berkopec's write-up quotes Heroku's test at 1.73× memory with default arenas against 0.87× with the limit, and he ends up recommending jemalloc outright for Puma and Sidekiq.

Go has a different problem: by default the heap is allowed to grow to twice the live set before the collector runs (GOGC=100), and the runtime has no idea there's a cgroup limit. Go 1.19 added GOMEMLIMIT, a soft limit that makes the GC work harder as the runtime's total memory approaches it. On the Mac (Apple M4, Go 1.25.4, since this is runtime behaviour, not kernel), a program holding 64 MB live and churning 2 GB of garbage:

Output
default:          peak held  151 MB  GCs   65
GOMEMLIMIT=90MiB: peak held   88 MB  GCs  262

Three runs each; the default ranged 151–159 MB and the limited one 88–90 MB. Put that program in a 128 MB container and the default run dies while the limited one doesn't. The GC guide says to leave "an additional 5-10% of headroom to account for memory sources the Go runtime is unaware of," and it caps GC CPU at roughly 50%, so a limit set too low slows you down about 2× instead of spinning forever.

The JVM went through this in public. Before container support, it sized its default max heap as a quarter of the host's RAM, so a 1 GB container on a 64 GB node got a 16 GB heap ceiling and a kernel OOM long before any OutOfMemoryError. Here's how it got fixed:

  1. JDK 9, backported to 8u131: experimental -XX:+UseCGroupMemoryLimitForHeap (JDK-8170888), off by default.
  2. JDK 10, backported to 8u191: -XX:+UseContainerSupport, on by default (JDK-8146115), plus MaxRAMPercentage (JDK-8186248). The experimental flag was deprecated, then removed.
  3. JDK 15, backported to 11.0.16 and 8u372: cgroup v2 support (JDK-8230305). An older JDK on a cgroup v2 node can quietly go back to sizing itself from host memory.

MaxRAMPercentage still defaults to 25.0 (gc_globals.hpp at jdk-21-ga), so a JVM in a 2 GB pod gets a 512 MB heap unless you tell it otherwise. And the heap isn't the whole process: metaspace, thread stacks, the code cache, direct buffers and glibc arenas all sit outside it.

So what does "memory usage" mean?

So the dashboard at 60% is usually plotting working set or RSS while memory.current runs at the limit, filled with things no process claims. Whether that ends in an OOM depends entirely on how much of it is clean page cache. Chapter 04 has the lower layer: why touching memory, not allocating it, is what makes these numbers move.

What to do on Monday

The machinery behind this story
More field notes