Field notes
The chapters are reference. These are one-sitting reads: a real outage, a bug you can reproduce on a laptop, a number everyone quotes that turns out to be wrong. Each one ends by pointing at the chapter that explains the machinery.
The 40-millisecond tax
Your request takes 41 ms on loopback, every time but the first. Two timers from the early 1980s are waiting for each other, and neither one is wrong.
Read itYou paid for four CPUs and got throttled at 40%
The dashboard says the pod is using 1.6 of its 4 cores. The kernel says it was throttled in 118 of the last 300 periods. Both are right. They just measure different things.
Read itThe dashboard said 60%. The OOM killer disagreed.
Your container has at least five different memory numbers, they disagree by hundreds of megabytes, and the kernel, the kubelet and your Grafana panel each trust a different one.
Read itfsyncgate: the fsync that returned 0 the second time
Postgres retried a failed fsync, the kernel said yes, and the data was already gone. I broke a disk on purpose to watch it happen on Linux 6.10.
Read itThe hour with 3,601 seconds in it
Twice in five years a single extra second took down large parts of the internet. Once it made Linux spin at 100% CPU, once it made Go measure a negative round trip. Both bugs are still easy to write.
Read itThe optimisation every database tells you to turn off
Transparent Huge Pages cut random-access latency by a quarter on my test box and made one page fault in a hundred take up to 22 milliseconds. The advice has flipped twice because both numbers are true.
Read it