KnowSys
Myth20 September 2026· ~15 min

The optimisation every database tells you to turn off

Transparent Huge Pages cut random-access latency by a quarter on my test box and made one page fault in a hundred take up to 22 milliseconds. The advice has flipped twice because both numbers are true.

linuxmemorydatabaseslatency

Every memory access your program makes goes through address translation. The CPU maps each virtual address to a physical one in fixed-size pages, 4 KiB on most Linux systems, and caches recent translations in a small buffer called the TLB. A big heap has millions of those pages, far more than the TLB can hold, so memory-heavy workloads spend real time looking translations up. Huge pages are 2 MiB instead, which cuts the number of translations by 512 times. Transparent Huge Pages (THP) is the kernel handing out huge pages on its own, without the application asking.

If you run Redis, MongoDB, Postgres, a JVM or anything else with a large heap, your hosts' THP setting affects you, whether or not anyone ever chose it. And the advice about it is unusually loud and unusually contradictory, which makes it worth understanding for yourself instead of copying a line from a runbook.

If you've run Redis 5 on a stock Linux box you've seen this line in the log, probably more than once, probably at 2am while looking for something else:

Output
WARNING you have Transparent Huge Pages (THP) support enabled in your kernel.
This will create latency and memory usage issues with Redis. To fix this issue
run the command 'echo never > /sys/kernel/mm/transparent_hugepage/enabled' as root

That's from src/server.c on the 5.0 branch. Redis isn't alone. Go looking and you'll find the same instruction from half the storage industry:

That's a strange consensus for a feature whose whole pitch is free speed. Then MongoDB 8.0 turned around and said the opposite: "ensure that THP is enabled before mongod starts" (docs, v8.0). Oracle's current guide now says to set it to madvise "and not disable Transparent HugePages as was recommended in prior releases" (26ai guide).

So which is it? I wanted to know what actually goes wrong, so I built each failure mode in a Linux VM and measured it. One of the famous explanations didn't reproduce at all.

What you're being offered

Chapter 04 covers TLB reach properly, so here's the one-line version. A TLB with around 1,500 entries covers about 6 MB of 4 KiB pages and about 3 GB of 2 MiB pages. A database doing random lookups across a multi-gigabyte heap misses the TLB on nearly every access with small pages, and each miss is a page-table walk.

Transparent Huge Pages (THP) is the kernel's attempt to hand out 2 MiB pages without the application asking. There's an older, explicit mechanism (hugetlbfs, what Postgres's huge_pages setting uses) where you reserve huge pages up front. THP is the automatic one, and it's the one everyone disables.

Three knobs decide how it behaves, all under /sys/kernel/mm/transparent_hugepage/:

Two machines, one kernel

Most of what follows ran in Docker Desktop, linuxkit 6.10.14, aarch64: a 10-vCPU VM with 7.8 GB of RAM and 4 KiB base pages, g++ 13.3 at -O2. Other jobs were running in that VM too. When it went away mid-way I reran the chase and fork tests on a shared 4-CPU container on the same 6.10.14 kernel, five interleaved runs each, and those medians are what I quote for them. I'd half expected THP to be compiled out of a small VM kernel. It wasn't:

Output
$ cat /sys/kernel/mm/transparent_hugepage/enabled
[always] madvise never
$ cat /sys/kernel/mm/transparent_hugepage/defrag
always defer defer+madvise [madvise] never
$ cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
511
$ cat /sys/kernel/mm/transparent_hugepage/hugepages-64kB/enabled
always inherit madvise [never]

Huge pages for everyone, and only processes that explicitly ask are allowed to stall for them. defrag=madvise is the upstream default; whether enabled starts as always or madvise is a build option, and distributions differ. Keep the 511 in mind. I didn't change any of these files, since other people were using the same kernel. Everything below is per-region madvise.

Reproducing the win

My test maps a 1 GiB buffer aligned to 2 MiB, applies MADV_HUGEPAGE or MADV_NOHUGEPAGE, and touches it. Then it does a dependent pointer chase in random order, one slot per 64-byte line, so every load waits on the previous one and the TLB sits on the critical path. Before trusting anything, it reads /proc/self/smaps to check the kernel really handed over huge pages.

1 GiB random chase, huge versus small pages
cpp
C++
char* p = map_aligned(1ul << 30);                 // mmap + round up to 2 MiB
madvise(p, bytes, huge ? MADV_HUGEPAGE : MADV_NOHUGEPAGE);
for (size_t off = 0; off < bytes; off += 4096) p[off] = 1;   // fault it in
anon_huge_kb(p);                                  // AnonHugePages from smaps
 
// one 8-byte slot per 64-byte line, linked into a single random cycle
for (size_t i = 0; i < lines; i++)
    v[order[i] * 8] = order[(i + 1) % lines] * 8;
for (size_t i = 0; i < 20'000'000; i++) idx = v[idx];   // timed
output
Output
huge    smaps: Rss 1048576 kB, AnonHugePages 1048576 kB
        random chase 136.5 ns/access   (median of 5; 125.9 – 151.4)
nohuge  smaps: Rss 1048576 kB, AnonHugePages       0 kB
        random chase 183.3 ns/access   (median of 5; 162.5 – 200.3)

Shared 4-CPU container. On the quieter first VM the best pair was 117.3 against 207.7 ns.

AnonHugePages 1048576 kB is the line to check on your own systems: the whole gigabyte was backed by 512 huge pages. And the win is real. A quarter off every random access on the shared box, closer to 40% on the quiet one. That's the number THP's defenders quote, and they aren't wrong.

The explanation everyone repeats

Here's why Redis says it hurts. Redis's latency guide says that "when a Linux kernel has transparent huge pages enabled, Redis incurs to a big latency penalty after the fork call is used." BGSAVE forks, the child writes the snapshot, and every page the parent modifies meanwhile gets copied (chapter 22 covers the memory side). With 2 MiB pages, the story goes, each copy is 512 times larger. I'd repeated that myself. It's in chapter 04 of this site.

So I measured it. After populating the buffer, my test forks, parks the child on a pipe, and has the parent write one byte into each 2 MiB region:

fork(), then one write per 2 MiB region in the parent
cpp
C++
pid_t pid = fork();                              // timed
if (pid == 0) { read(pipefd[0], &c, 1); _exit(0); }   // child just holds the pages
 
for (size_t off = 0; off < bytes; off += 2 << 20) p[off + 64] = 2;   // 512 writes, each timed
// then: AnonHugePages for the region, and Private_Dirty from smaps_rollup
output
Output
huge    fork() 0.51 ms
        COW, 1 write per 2 MiB: 512 writes, median 4.4 us, max 69.9 us
        smaps: AnonHugePages 0 kB
        rollup Private_Dirty: 2120 kB
nohuge  fork() 12.93 ms
        COW, 1 write per 2 MiB: 512 writes, median 2.1 us

First VM, one run. Shared-container medians of five: fork() 0.38 ms against 3.04 ms, and 3.9 µs against 1.2 µs per write.

Look at Private_Dirty. If each write had copied a 2 MiB page, the parent would own a gigabyte of fresh private memory. It owns 2,120 kB: 512 copies of 4 KiB plus change. Meanwhile AnonHugePages went from a full gigabyte to zero. So the kernel didn't copy the huge pages. It broke them.

Here's that path in the kernel I ran, trimmed:

mm/huge_memory.c
linux @ v6.10 ↗
C
vm_fault_t do_huge_pmd_wp_page(struct vm_fault *vmf)
{
	...
	/* Early check when only holding the PT lock. */
	if (PageAnonExclusive(page))
		goto reuse;            // nobody else maps it: just make it writable
	...
	if (folio_ref_count(folio) >
			1 + folio_test_swapcache(folio) * folio_nr_pages(folio))
		goto unlock_fallback;  // the fork child still maps it
	...
unlock_fallback:
	folio_unlock(folio);
	spin_unlock(vmf->ptl);
fallback:
	__split_huge_pmd(vma, vmf->pmd, vmf->address, false, NULL);
	return VM_FAULT_FALLBACK;  // retry as an ordinary 4 KiB COW fault
}

There's no branch left that allocates a fresh 2 MiB page and copies into it. Two things follow, and neither is the folklore version:

Snapshots do have a quieter cost: they shred your huge pages. After one BGSAVE on a write-heavy instance, most of the heap is back on 4 KiB pages until khugepaged gets around to rebuilding it. Which brings us to the costs that are real.

What actually hurts: stalling at fault time

Remember defrag. When a process faults on a huge-page-eligible region and the buddy allocator has no free 2 MiB block, the kernel has to decide whether to make one. Making one means direct compaction: migrating other pages out of the way, synchronously, inside your page fault. Here are the options, quoted from the kernel docs where the wording matters:

always used to be the default. Mel Gorman changed it to madvise in Linux 4.6, in commit 444eb2a449ef, whose message says "It's been years and it's time to throw in the towel." David Rientjes of Google added defer+madvise in 4.11. A lot of the 2012–2016 horror stories come from always, on RHEL 6 and 7 kernels that predate both changes.

I measured it directly. A second test faults a 1 GiB region one 2 MiB chunk at a time and times each chunk. Three flavours: MADV_HUGEPAGE (allowed to compact, under the default defrag=madvise), no advice at all (THP allowed but not allowed to stall), and MADV_NOHUGEPAGE. Then I ran it again after fragmenting memory: fill 1.5 GB with small pages and free every other one, leaving holes everywhere.

RegionMemory statep50 per 2 MiBp99maxGot huge pages
MADV_HUGEPAGEfresh29–31 µs2.7–11.7 ms11–15 ms100%
no advicefresh26–31 µs0.29–0.41 ms0.4–2.0 ms98–99%
MADV_NOHUGEPAGEfresh174–475 µs0.36–2.8 ms0.5–3.7 ms0%
MADV_HUGEPAGEfragmented267–663 µs14.9–22.1 ms18–35 ms80–84%
no advicefragmented180–413 µs0.7–5.4 ms1.5–9.6 ms37–42%
MADV_NOHUGEPAGEfragmented194–195 µs0.65–0.74 ms0.9 ms0%

Three runs on fresh memory and two fragmented, on the first VM. A small-page chunk pays 512 separate faults, so that row's median is worst.

There's the whole trade in one table. At the median, huge pages are six to fifteen times cheaper to fault in. At the tail, on fragmented memory, a region that asked for them nicely waits 15 to 22 milliseconds at p99 and 35 at worst. During a single fresh-memory run /proc/vmstat counted 100 new compact_stall events. One fault in a hundred took longer than forking the whole gigabyte did with small pages. You won't find it in a CPU profile either, because it's sitting inside a store instruction in your malloc.

Now compare the first two rows. On fresh memory, the no-advice region fell back to small pages 1–2% of the time, and the MADV_HUGEPAGE region's p99 is bad. I think those are the same 1–2% of faults: where one gave up, the other waited for compaction. That's defrag=madvise doing exactly what the docs say, and it's also why the no-advice region was only 37–42% huge once memory was fragmented.

What actually hurts: memory you didn't ask for

RSS is the other real cost. A huge page is all-or-nothing: touch one byte and you've got 2 MiB resident. A third test touches one byte in each 2 MiB region of a 256 MiB mapping:

Output
NOHUGEPAGE, 1 byte per 2 MiB       Rss     512 kB   AnonHugePages       0 kB
HUGEPAGE,   1 byte per 2 MiB       Rss  262144 kB   AnonHugePages  262144 kB

Same 128 bytes of data, 512 times the memory. Real allocators aren't quite that sparse, but they're sparser than you'd think. jemalloc returns memory to the kernel with madvise(MADV_DONTNEED) on ranges usually much smaller than 2 MiB. DigitalOcean found in 2015 that under THP those ranges often couldn't be reclaimed "because the entire page would have to be unneeded" (DigitalOcean engineering). jemalloc's author wrote that its purging "wreaks havoc on Linux transparent huge pages" (issue #243), and a 2019 report measured 250–283 MB RSS with THP always against 12–17 MB with it off, for the same 100-thread program.

khugepaged makes this worse slowly, and slowly is the worst way. Its max_ptes_none is how many empty 4 KiB slots it'll fill with zeroes when collapsing a 2 MiB range. In mm/khugepaged.c the default is HPAGE_PMD_NR - 1, so 511: it'll build a huge page around a single live 4 KiB page. I touched one page in eight of a fresh 256 MiB region, then called madvise(MADV_HUGEPAGE) on it and watched:

Output
1 page in 8, before madvise        Rss   32768 kB   AnonHugePages       0 kB
  after MADV_HUGEPAGE, t=40s       Rss   32768 kB   AnonHugePages       0 kB
  after MADV_HUGEPAGE, t=50s       Rss   36352 kB   AnonHugePages    4096 kB
  after MADV_HUGEPAGE, t=100s      Rss  108032 kB   AnonHugePages   86016 kB
  after MADV_HUGEPAGE, t=180s      Rss  222720 kB   AnonHugePages  217088 kB

Nothing for 40 seconds, then RSS climbed about 14 MiB every ten seconds. With pages_to_scan at 4096 and a 10-second sleep, khugepaged covers 16 MiB per pass across the whole machine, so that's about the rate you'd predict. My process didn't allocate a byte in those three minutes, and its footprint went up almost sevenfold. Now picture that inside a container with a memory limit. That's Couchbase's "possibly lead to an OOM kill."

Why the advice changed

None of these costs went away. What changed is who's in charge of the tradeoff. In the old arrangement, a kernel guessed on every fault whether memory was worth backing with 2 MiB, for an allocator with no idea huge pages existed. In the new one, an allocator packs objects to fill huge pages, and a kernel only stalls when asked.

Google did the allocator half first. Their OSDI 2021 paper, Beyond malloc efficiency to fleet efficiency, describes Temeraire, a huge-page-aware backend for TCMalloc. Across eight application studies it improved "requests-per-second (RPS) by 7.7% and reducing RAM usage 2.4%", and at fleet scale it gave "6% fewer TLB miss stalls, and 26% reduction in memory wasted due to fragmentation." Look at that last number. A huge-page-aware allocator fragments less, because it knows which 2 MiB ranges it can hand back whole.

That's the TCMalloc MongoDB 8.0 moved to, and it's why its advice reversed. Its 8.0 docs ask for a specific combination:

It applies only on x86_64 and ARM64 with a kernel new enough for rseq (4.18+), and "If you are using MongoDB 7.0 or earlier, disable THP." Read that list again next to the measurements above. Each line switches off one of the failure modes I reproduced.

Meta came at it from the kernel side. Their 2024 patch series splitting underused THPs opens by saying "Meta uses madvise in production as the current THP=always policy vastly overprovisions THPs in sparsely accessed memory areas." On CPU-bound services, always bought 1.8% performance for 7.7% more memory; with the new shrinker, about the same gain (1.7%) cost only 2.4%. It landed in Linux 6.12 as transparent_hugepage/shrink_underused. That's newer than the 6.10 kernel I tested, so I haven't run it. jemalloc has an experimental huge-page-aware allocator too (src/hpa.c in 5.3.0), but it's off by default and I wouldn't bet a database on it yet.

And the pages got smaller. Ryan Roberts's multi-size THP landed in Linux 6.8 (LWN) and lets anonymous memory use folios between 4 KiB and 2 MiB. You can see the sysfs directories in my transcript above: 16 kB through 1024 kB, all never by default. A 64 KiB page probably wastes far less on a sparse heap and needs no compaction to find, but those switches are global, and I wasn't going to flip them on a kernel other people were benchmarking on. So I can't tell you what they'd do to my table.

Redis moved too, more quietly. Since 6.2 it ships disable-thp yes and calls prctl(PR_SET_THP_DISABLE) on itself when the system is set to always (6.2 release notes). Its current startup check recommends madvise instead of never. A modern Redis already opts itself out, so you don't need to turn THP off for the whole box to protect it.

Even for plain JVMs it cuts both ways. Alexandr Nikitin measured THP improving p95 by about 0.8 ms on a 200 GB-heap service in 2017. LinkedIn in 2013 turned it off after the kernel split a 5 GB heap's huge pages to migrate them between NUMA nodes. Same feature, different workloads, opposite answers.

What I got wrong, and what I still can't explain

My first pass had an anomaly I was ready to write up. After the fork, the parent also wrote into every 4 KiB page, and on the first VM the pages that had been huge came out cheaper to COW (1.2–3.2 µs) than pages that never were (2.6–4.0 µs). I had a theory about physical contiguity and struct page locality. On the quieter shared box the gap vanished: 1.14 µs against 1.06 µs, medians of five. It was probably just the other jobs on the first VM.

What I can't explain is one fragmented run of the no-advice region. Under defrag=madvise that region isn't allowed to compact, yet its p99 hit 5.4 ms, against 0.7 ms on the run beside it. Maybe it's direct reclaim on the small-page fallback path, or maybe it's another tenant stealing the CPU. I didn't record /proc/vmstat around that run, and the shared box caps buffers at 1 GB, so I couldn't recreate the 1.5 GB fragmentation to find out. If you can, capture compact_stall, allocstall_movable and pgscan_direct before and after; that should settle it.

What to do on Monday

Start by finding out what you have. These are safe to run on any box:

Shell
cat /sys/kernel/mm/transparent_hugepage/enabled /sys/kernel/mm/transparent_hugepage/defrag
cat /sys/kernel/mm/transparent_hugepage/khugepaged/max_ptes_none
grep AnonHugePages /proc/$PID/smaps_rollup           # is your process actually using them?
grep -E 'thp_fault_alloc|thp_fault_fallback|compact_stall' /proc/vmstat
uname -r                                              # below 5.8, the 2 MiB COW copy is real

Then decide:

Alert on compact_stall rising in /proc/vmstat. It's the counter that moves when a page fault started waiting on the kernel, and in my runs it was the only sign those 35 ms faults had happened at all. And in design review, ask two questions: does anything here fork() a large heap, and which kernel is it on? Those two answers are most of the decision.

The machinery behind this story
More field notes