KnowSys

Field notes

The chapters are reference. These are one-sitting reads: a real outage, a bug you can reproduce on a laptop, a number everyone quotes that turns out to be wrong. Each one ends by pointing at the chapter that explains the machinery.

cgroupsclocksdatabasesdistributed-systemsdurabilityfilesystemskuberneteslatencylinuxmemorypostgresschedulingtcptime
Reproduction25 September 2026· ~13 min

The 40-millisecond tax

Your request takes 41 ms on loopback, every time but the first. Two timers from the early 1980s are waiting for each other, and neither one is wrong.

Read it
Incident24 September 2026· ~14 min

You paid for four CPUs and got throttled at 40%

The dashboard says the pod is using 1.6 of its 4 cores. The kernel says it was throttled in 118 of the last 300 periods. Both are right. They just measure different things.

Read it
Field guide23 September 2026· ~16 min

The dashboard said 60%. The OOM killer disagreed.

Your container has at least five different memory numbers, they disagree by hundreds of megabytes, and the kernel, the kubelet and your Grafana panel each trust a different one.

Read it
Incident22 September 2026· ~15 min

fsyncgate: the fsync that returned 0 the second time

Postgres retried a failed fsync, the kernel said yes, and the data was already gone. I broke a disk on purpose to watch it happen on Linux 6.10.

Read it
Incident21 September 2026· ~14 min

The hour with 3,601 seconds in it

Twice in five years a single extra second took down large parts of the internet. Once it made Linux spin at 100% CPU, once it made Go measure a negative round trip. Both bugs are still easy to write.

Read it
Myth20 September 2026· ~15 min

The optimisation every database tells you to turn off

Transparent Huge Pages cut random-access latency by a quarter on my test box and made one page fault in a hundred take up to 22 milliseconds. The advice has flipped twice because both numbers are true.

Read it