KnowSys
Reproduction25 September 2026· ~13 min

The 40-millisecond tax

Your request takes 41 ms on loopback, every time but the first. Two timers from the early 1980s are waiting for each other, and neither one is wrong.

tcplinuxlatency

Nearly every request your service makes rides on TCP: the HTTP call to the next service over, the query to Postgres, the GET to Redis. You rarely think about it, because on a healthy network it adds a fraction of a millisecond and gets out of the way. Underneath, the kernel is making small decisions for you on every connection, like when to send a packet and when to send an acknowledgement back, and its defaults were picked for networks that looked nothing like a data centre.

Usually those defaults are invisible. Once in a while two of them line up and put a fixed delay on every request, and faster hardware won't remove it, because nothing is slow. The kernel is waiting on purpose. If you run anything that does request and reply over a long-lived connection, you'll want to recognise it on sight, since it passes for a performance problem and isn't one.

Here's the output that started this. It's a client and a server in one C++ program, talking over a single persistent TCP connection on 127.0.0.1. The client sends a request, waits for a 128-byte reply, and does it again. No disk, no database, no network card. Numbers are microseconds per round trip, and the last column is the first twelve requests in the order they happened:

C++
default  n=200  p50  41103.0 us  p90  42293.3 us  p99  44192.9 us | first 12: 29 41232 40763 41708 40238 42087 ...

The first request took 29 µs. Every request after it took about 41,000. Not a spread from 30 µs to 80 ms, the way a noisy system looks. A cliff, then a flat line at 41 ms, three runs in a row, to within a millisecond.

If you've been paged for this, you've probably seen it as a latency histogram with a spike parked at 40 ms (or 200 ms, on some systems), and a p50 that makes no sense next to a ping time of a fraction of a millisecond. Before getting to what it is, it's worth walking through what it isn't, because the wrong theories are the ones everyone reaches for first. I reached for two of them.

Theory one: it's GC, or the scheduler

A flat 41 ms smells like a pause. In a Java or Go service the first suspect is a collector, and in anything else it's a thread that got descheduled.

Neither fits. This program is C++ with no allocator in the loop, so there's nothing to collect. And pauses don't behave like this. A GC pause lands on some requests and not others, and its length depends on the heap. Scheduler delays are jittery, and at 41 ms you'd be looking at a machine so overloaded that top would tell you before any histogram did. This is every request, the same length each time. Something is counting to 40.

Theory two: it's DNS, or connection setup

This is the one I half-believed for a minute, because the first request is the odd one out. Maybe the first request is special because it's the only one that got something for free?

That has it backwards, though. DNS, TCP handshakes and TLS all make the first request slow and the rest fast. Here the connection is set up once, before the timer starts, and the first request is the fast one. Whatever this is, it's a property of a connection that has been alive for a little while.

Theory three: it's the network

Is it the network? On loopback there isn't one, honestly. ss -tin on the live connection reports a smoothed RTT of 0.03 ms on the server side and a minimum RTT of 0.003 ms on the client. A round trip on this path costs microseconds. You can't get 41 ms out of it by waiting on wires.

What you can get is 41 ms of waiting on a timer. And once the question is "which timer in the TCP stack is 40 ms long?", the list is short.

The shape of the request

My client doesn't send its request in one piece. It does what almost every HTTP client with a naive send path does: headers in one write(), body in a second one, then a read() for the response.

C++
write(c, hdr, 64);     // write #1: headers
write(c, body, 256);   // write #2: body
read_full(c, resp, 128);

On the other end, the server reads until it has all 320 bytes of the request, then sends one 128-byte response. That's all it does. This is the write-write-read pattern, and it's the entire bug.

Two mechanisms meet here, and each is reasonable alone:

Now run the pattern through both. Headers go out at once, because nothing is in flight. Then the body: it's small, and the header hasn't been ACKed, so Nagle holds it. Meanwhile the server has 64 bytes of a 320-byte request, so it has nothing to reply with. Its stack is betting that a reply is coming, so it holds the ACK. Each side is waiting for the other, and the only thing that breaks the tie is the delayed-ACK timer expiring.

clientkernel (loopback)server~41 ms: both sides waitingheaders, 64 Bsent: nothing in flightbody, 256 BNagle: held, headers unACKeddelayed ACKno reply yet to ride onpure ACKdelack timer firesbody, 256 Breleasedresponse, 128 B
Every request after the first. The client's body sits in its own send queue for the entire delayed-ACK timeout, waiting for an ACK the server is deliberately holding back.

Watching it happen

I ran the reproduction in docker run --privileged ubuntu:24.04 on Docker Desktop (linuxkit 6.10.14 on aarch64, an Apple M4 underneath), g++ 13.3 -O2, with tcpdump -i lo -ttt running alongside. -ttt prints the time since the previous packet, and that's exactly the view you want:

C++
.000015  client > server  [P.] seq 1:65      length 64    # request 1: headers
.000002  server > client  [.]  ack 65        length 0     #   ACKed at once
.000004  client > server  [P.] seq 65:321    length 256   #   body
.000027  server > client  [P.] seq 1:129     length 128   #   response
.000006  client > server  [P.] seq 321:385   length 64    # request 2: headers
.042416  server > client  [.]  ack 385       length 0     #   42 ms later: pure ACK
.000016  client > server  [P.] seq 385:641   length 256   #   body, finally
.000018  server > client  [P.] seq 129:257   length 128   #   response
.000042  client > server  [P.] seq 641:705   length 64    # request 3: headers
.042573  server > client  [.]  ack 705       length 0     #   and again

(I replaced the addresses with client and server, dropped the TCP options and added the comments. No packets were removed.)

There it is on the wire. Request one's headers get ACKed in 2 µs. Request two's headers sit unacknowledged for 42.4 ms, the body doesn't leave the client until the ACK lands, and the pattern repeats for every request after that. On the same connection ss -tin shows ato:40. That's the kernel's current acknowledgement timeout for this socket, in milliseconds.

So what is the Linux delayed-ACK timer, exactly?

It isn't a sysctl. It's a pair of constants in include/net/tcp.h, in jiffies:

include/net/tcp.h
linux @ v6.10 ↗
C
#define TCP_DELACK_MAX	((unsigned)(HZ/5))	/* maximal time to delay before sending an ACK */
static_assert((1 << ATO_BITS) > TCP_DELACK_MAX);
 
#if HZ >= 100
#define TCP_DELACK_MIN	((unsigned)(HZ/25))	/* minimal time to delay before sending an ACK */
#define TCP_ATO_MIN	((unsigned)(HZ/25))
#else
#define TCP_DELACK_MIN	4U
#define TCP_ATO_MIN	4U
#endif

HZ/25 is 40 ms and HZ/5 is 200 ms, whatever CONFIG_HZ is (this kernel runs CONFIG_HZ=1000). Each socket carries its own ato. It starts at TCP_ATO_MIN and can grow when the kernel misses a chance to piggyback. So on Linux the delay you pay is somewhere from 40 to 200 ms. My loopback case sits on the floor. GoCardless hit the ceiling: every internal POST in their Ruby stack carried about 200 ms of it until they upgraded HAProxy.

Why is the loopback case stuck at 40 when the RTT is 30 µs? You'd hope the kernel would scale the delay down on a fast path, and it does try. Just not here:

net/ipv4/tcp_output.c — tcp_send_delayed_ack()
linux @ v6.10 ↗
C
	int ato = icsk->icsk_ack.ato;
 
	if (ato > TCP_DELACK_MIN) {
		...
		/* If some rtt estimate is known, use it to bound delayed ack. */
		if (tp->srtt_us) {
			int rtt = max_t(int, usecs_to_jiffies(tp->srtt_us >> 3),
					TCP_DELACK_MIN);
			if (rtt < max_ato)
				max_ato = rtt;
		}
		ato = min(ato, max_ato);
	}
 
	ato = min_t(u32, ato, tcp_delack_max(sk));
	timeout = jiffies + ato;

That RTT bound only kicks in when ato is above the minimum, and even then it's clamped to be no lower than TCP_DELACK_MIN. At the floor, 40 ms is the answer regardless of how fast the path is. The extra millisecond in my 41 ms is probably timer slop plus the actual work, though I didn't try to break it down.

On the sender, the code is just as short. Linux implements Minshall's variant of Nagle: hold a partial segment only if an earlier partial segment is still unacked.

net/ipv4/tcp_output.c — tcp_nagle_check()
linux @ v6.10 ↗
C
static bool tcp_minshall_check(const struct tcp_sock *tp)
{
	return after(tp->snd_sml, tp->snd_una) &&
		!after(tp->snd_sml, tp->snd_nxt);
}
 
static bool tcp_nagle_check(bool partial, const struct tcp_sock *tp,
			    int nonagle)
{
	return partial &&
		((nonagle & TCP_NAGLE_CORK) ||
		 (!nonagle && tp->packets_out && tcp_minshall_check(tp)));
}

On loopback the MSS is 32,768 bytes (it's in the ss output), so a 64-byte header is as partial as a segment gets. snd_sml is the end of the last small segment sent. While it's past snd_una, the next small write waits.

Why the first request was fast

This is the part that took me longest, and I think it's the most interesting piece of the story. Linux doesn't delay ACKs on every connection. It delays them once it has decided the connection is interactive. It calls that pingpong mode, and the decision happens when the socket sends data:

C
/* net/ipv4/tcp_output.c, tcp_event_data_sent() */
	if ((u32)(now - icsk->icsk_ack.lrcvtime) < icsk->icsk_ack.ato)
		inet_csk_inc_pingpong_cnt(sk);

"If I'm sending data back within ato of receiving some, this looks like request and response, so delaying ACKs will pay off." On request one the server hasn't replied to anything yet, so it's still in quick-ACK mode and ACKs the headers at once. Its reply to request one bumps the pingpong count past the threshold, and from request two onwards every ACK is delayed.

Then it gets stranger. When the delayed-ACK timer fires in pingpong mode, the kernel concludes it bet wrong and backs off:

net/ipv4/tcp_timer.c — tcp_delack_timer_handler()
linux @ v6.10 ↗
C
	if (inet_csk_ack_scheduled(sk)) {
		if (!inet_csk_in_pingpong_mode(sk)) {
			/* Delayed ACK missed: inflate ATO. */
			icsk->icsk_ack.ato = min_t(u32, icsk->icsk_ack.ato << 1, icsk->icsk_rto);
		} else {
			/* Delayed ACK missed: leave pingpong mode and
			 * deflate ATO.
			 */
			inet_csk_exit_pingpong_mode(sk);
			icsk->icsk_ack.ato      = TCP_ATO_MIN;
		}
		tcp_send_ack(sk);

So it does learn. It leaves pingpong mode and resets ato to 40 ms. And then the server sends its response within ato of receiving the body, so the count goes up again and pingpong mode is back before the next request arrives. So the kernel learns its lesson and unlearns it about 20 µs later, once per request.

How fast it re-learns is a sysctl, net.ipv4.tcp_pingpong_thresh, default 1 (ip-sysctl docs). I changed it inside the container and reran the same client:

tcp_pingpong_threshp50first 12 requests (µs)
141,004 µs21 40356 41219 40806 40979 40990 ...
23,008 µs29 34 40132 31 41017 37 40952 41 ...
392 µs73 94 61 40174 50 47 42135 ...
552 µs49 57 41 73 90 40664 36 36 40 49 40818 ...

With a threshold of n, roughly one request in n pays the 40 ms. With 2 you get a sawtooth, and a p50 of 3 ms that corresponds to no request that ever happened. That's a nice trap for anyone who only looks at percentiles. Raising the threshold makes the median look fine and leaves the tax on the tail. That's the textbook way for contention to hide in a p99.

The fixes, measured

Same program, same container, 500 requests per mode, three runs. The microsecond numbers bounce around by a factor of three between runs. That's Docker Desktop's VM, more or less, doing what it does to anything this short. Milliseconds don't move at all.

One process, two threads, one persistent loopback connection
cpp
C++
// server thread
for (;;) {
    read_full(s, req, 64 + 256, mode == "quickack-loop"); // whole request
    if (mode == "quickack")                               // once, after reading
        setsockopt(s, IPPROTO_TCP, TCP_QUICKACK, &one, sizeof one);
    write(s, resp, 128);
}
 
// read_full: with quick=true, re-arm TCP_QUICKACK before every read() call
while (n) {
    if (quick) setsockopt(fd, IPPROTO_TCP, TCP_QUICKACK, &one, sizeof one);
    ssize_t r = read(fd, p, n); p += r; n -= r;
}
 
// client
if (mode == "nodelay") setsockopt(c, IPPROTO_TCP, TCP_NODELAY, &one, sizeof one);
for (int i = 0; i < iters; i++) {
    if (mode == "writev") {
        iovec v[2] = {{hdr, 64}, {body, 256}};
        writev(c, v, 2);                 // one syscall, one segment
    } else {
        write(c, hdr, 64);
        write(c, body, 256);
    }
    read_full(c, resp, 128);
}
output
modep50, three runsp99, three runs
default41.0 / 41.0 / 41.0 ms43.6 / 50.7 / 51.4 ms
TCP_NODELAY on the client20 / 28 / 31 µs70 / 753 / 88 µs
one writev7 / 29 / 16 µs31 / 35 / 61 µs
TCP_QUICKACK once, after the read41.0 / 41.0 / 41.0 ms45.6 / 45.9 / 45.1 ms
TCP_QUICKACK before every read()17 / 6 / 32 µs39 / 17 / 63 µs

Three things fix it and one doesn't. And the one that doesn't is what Nagle himself recommends, set in the obvious place.

TCP_NODELAY removes the sender's half of the deadlock. writev removes the pattern itself: one syscall and one segment, so there's no second small write for Nagle to hold. It's also one fewer kernel crossing per request, and that matters on its own when a syscall costs hundreds of nanoseconds. Of the three, it's the one I'd reach for first.

TCP_QUICKACK is the subtle one. tcp(7) says so, if you read to the end of the entry: "This flag is not permanent, it only enables a switch to or from quickack mode. Subsequent operation of the TCP protocol will once again enter/leave quickack mode." Set it after the read, and by the time the next request's headers arrive, the reply you just sent has put the socket back into pingpong mode. Set it before each read, when you're holding a partial request, and it works. That's exactly what HAProxy 1.4.19 did for GoCardless: when a packet holds only the first part of a POST, it enables TCP_QUICKACK and ACKs immediately.

What John Nagle thinks of this

Julia Evans wrote about this exact bug on 21 November 2015, in Why you should understand (a little) about TCP. Someone at work was publishing messages to NSQ on localhost and each one took 40 ms. Ruby's Net::HTTP split each POST across two packets, headers then body, with Nagle on. She set TCP_NODELAY, and "all of the 40ms delays instantly disappeared."

The same day, on the Hacker News thread for her post, John Nagle showed up (he posts as Animats):

That still irks me. The real problem is not tinygram prevention. It's ACK delays, and that stupid fixed timer. They both went into TCP around the same time, but independently. I did tinygram prevention (the Nagle algorithm) and Berkeley did delayed ACKs, both in the early 1980s. The combination of the two is awful.

He explains that delayed ACKs made sense for Berkeley's Telnet traffic from terminal rooms, that "a delayed ACK is a bet that the other end will reply to what you just sent almost immediately," and that outside some RPC protocols the bet keeps losing. His short version is "set TCP_QUICKACK." My table above is a footnote to that: set it, but on Linux, set it again before every read.

He's also careful about the other fix. With TCP_NODELAY, a loop of one-byte write() calls becomes one packet per byte, "a factor of 40" more traffic. So TCP_NODELAY doesn't excuse small writes; it just stops the kernel from papering over them. Marc Brooker's 2024 post, It's always TCP_NODELAY. Every damn time., takes the opposite default: check TCP_NODELAY first, and turn Nagle off for modern systems.

Who already turned it off

A lot of the software you run has made this call for you. That's probably why you mostly meet the bug in older or hand-rolled clients:

If your stack isn't on that list, it's worth ten minutes with strace -e setsockopt to find out what it does.

And on macOS?

I ran the same program natively on the Mac (Apple M4, macOS 26.5.1, clang -O2), and it didn't reproduce:

C++
default   n=1000  p50  35.9 us  p90  52.7 us  p99  64.2 us  max  94.7 us
nodelay   n=200   p50  40.8 us  p90  43.1 us  p99  47.5 us  max  49.8 us
writev    n=200   p50  37.7 us  p90  41.1 us  p99  49.7 us  max  53.6 us

No cliff. XNU does delay ACKs (net.inet.tcp.delayed_ack is 3 here), and its timer is longer than Linux's: tcp_delack = TCP_RETRANSHZ / 10, with a comment saying it "will fire some where between 100 and 200 ms" (bsd/netinet/tcp_timer.c). So if it were holding these ACKs, I'd have seen 100 ms, not 36 µs.

XNU's decision lives in tcp_cc_delay_ack() in bsd/netinet/tcp_cc.c. Mode 3 delays only when seven conditions all hold. One is that the receive buffer holds more than its low-water mark. My guess is that a server thread already blocked in read() empties the buffer fast enough that the condition fails, so the ACK goes straight out. I haven't confirmed it. Packet capture on macOS needs root for /dev/bpf. I didn't have that for this run, and the per-protocol counters in netstat -s read zero for delayed ACKs before and after. So this stays open: Linux on loopback reproduces every time, macOS on loopback doesn't, and I can point at the code but not at the packet that proves why.

There's a second loose end in the Linux capture. After the timer-driven ACK releases the body, the server sends another pure ACK for the body about 20 µs before its response, where I expected the ACK to ride on the response. My best reading is that leaving pingpong mode makes the ACK go out as soon as the application drains the buffer, before the write() gets there. I haven't traced it.

What to do on Monday

In design review, the question is just: "how many writes does one request take, and is Nagle on?" If the answer is "more than one" and "yes", you've found your 40 ms before anyone gets paged for it.

The machinery behind this story
More field notes