Your program has a line to record, buy milk, and it calls write() to add that line to a log file. It's one line of code, and it returns almost at once. Run it again for the next request, and again for the one after, and a busy service ends up doing it hundreds of thousands of times a second.
Now look at what that line asks for. A log file lives on a disk, and your program can't reach the disk. The processor won't let ordinary program code touch disks, network cards, or the memory of other programs. The code that is allowed to is the kernel, the core of the operating system, which runs with powers your program doesn't have. So when your program calls write(), someone else does the work: the request travels to the kernel, the kernel decides whether to honour it, does the writing, and reports back.
Each of those trips is quick, and a million of them are not. This chapter follows that one write() of buy milk through the trip and asks three things: what happens on the way, what does it cost, and how do fast programs avoid paying it so often? We start by timing it, because the answer to the second question is a surprise.
01Timing a million small writes
1.1A million tiny trips against a few big ones
Here's the smallest experiment that shows the whole chapter. It sends about a megabyte to a place that throws everything away, once as a million separate one-byte calls, and once as a few hundred calls of 4 KB each (4,096 bytes).
A few pieces of the script need introducing. When a program opens a file, the kernel hands back a file descriptor, a small number the program passes to every later call to say which file it means. os.open returns one here, and os.write(fd, data) is Python's thin wrapper around the kernel's write call. The file being opened is /dev/null, a special file that accepts any bytes you write to it and discards them. Using it means no disk is involved, so any time the script spends belongs to the calls themselves. time.perf_counter() reads a stopwatch in seconds, and the script reads it before and after each loop.
Save it as cross.py and run it with python3 cross.py. Any Mac or Linux machine will do.
import os, time
fd = os.open("/dev/null", os.O_WRONLY) # a file that throws away what you write
n = 1_000_000
t = time.perf_counter()
for _ in range(n):
os.write(fd, b"x") # one byte per call
a = time.perf_counter() - t
block = b"x" * 4096
t = time.perf_counter()
for _ in range(n // 4096 + 1):
os.write(fd, block) # 4 KB per call
b = time.perf_counter() - t
print(f"1,000,000 writes of 1 byte : {a*1000:7.1f} ms")
print(f" 245 writes of 4 KB : {b*1000:7.1f} ms")1,000,000 writes of 1 byte : 378.1 ms
245 writes of 4 KB : 0.1 msBoth versions send the same megabyte. The first makes a million calls and takes over a third of a second. The second makes 245 calls and finishes in about a tenth of a millisecond, thousands of times sooner. The exact times change from machine to machine and from run to run, but the second line always comes out thousands of times smaller than the first.
1.2What the numbers say
Nothing was stored anywhere, since /dev/null discards its input, so the time went into the calls themselves. Divide 378 milliseconds by a million and each call cost about 380 nanoseconds, where a nanosecond is a billionth of a second. That figure includes Python's own loop, so the trip itself is a little cheaper: section 3 puts it at about 350.
Three hundred and fifty nanoseconds is nothing once. A million times, it's a third of a second, and fast programs are the ones that make fewer trips with more in each. But a call that does no work at all still spends hundreds of nanoseconds, and that needs explaining. To see where the time goes, we first have to ask why there's a border to cross.
02Why there's a border at all
2.1Two privilege modes
?Why can't your program just talk to the disk itself?
Because then every program could. A bug in your text editor could scribble over the filesystem, and any program could read another program's memory or reprogram the network card to watch other people's traffic. Something has to stand between programs and the hardware.
The processor enforces this itself. It runs in one of (at least) two modes. In user mode, where your code runs, the instructions that talk to hardware or change which memory belongs to which program are forbidden. If your program tries one, the processor stops it with a fault and hands control to the kernel, which usually ends the program. In kernel mode those instructions work, and only the kernel runs there. The mode a piece of code runs in is its privilege level. (ARM processors, such as the ones in recent Macs, call these levels EL0 for user code and EL1 for the kernel.)

2.2The one door in
That leaves a puzzle. If your program can't touch the disk, how does it ever read a file?
Think of a bank vault with a teller window. You can't walk into the vault, and the bank doesn't want other customers walking in either. So you fill out a slip, hand it through the window, and wait while the teller checks it, fetches what you asked for, and passes it back. Your program asks the kernel for things the same way, and the processor provides the window.
?Why not let programs call into the kernel like any other function?
Because a program that could jump to any address in the kernel could jump past the line that checks its permissions. So the processor offers exactly one way in: a special instruction (syscall on x86-64 chips, svc on ARM chips such as Apple's), which we'll call the trap instruction. Before running it, your program puts a number in a register, one of the handful of tiny storage slots inside the processor, to say which service it wants: read, write, or openat (open a file). The arguments go in other registers. The trap instruction switches the processor to kernel mode and jumps to a single entry point that the kernel chose when the machine booted.
That whole request through the door is a system call, or syscall. It's a controlled transfer into code you can't see, running at a privilege level you don't have. The kernel promises to check every argument, do the thing, and hand control back, or to fail with an error code and leave the system intact.

2.3Why the kernel copies your arguments
The kernel can't trust anything that comes through the door, and the arguments are the first place that shows.
?Why doesn't the kernel just use the pointer you gave it?
A pointer is an address in memory where data lives, and write takes one to say where your bytes are. A pointer you hand over might point into the kernel's own memory, which would trick the kernel into reading its secrets for you. And your program may have several threads, separate flows of execution that share memory, so another thread could change the buffer while the kernel is halfway through reading it.
So the kernel checks every pointer and length, then copies the bytes into memory of its own and works from the copy. That copy is a real cost, which is part of why large transfers behave differently from small ones.
That's the design. Now we can follow our one write() of buy milk through it, step by step, and put a price on the trip.
03Inside one write()
3.1One write(), across the border and back
Our program calls write(fd, buf, 9). The first argument is the descriptor of the log file; say the program opened the log earlier and the kernel handed back descriptor 3. buf is a pointer to the line itself, the eight characters of buy milk followed by the newline that ends the line, and 9 is how many bytes that makes. The function your code calls isn't the kernel's. It lives in libc, the standard C library that nearly every program on a Unix-like system links in, and it's a small wrapper whose only job is to set up the trap.

The kernel also needs two things of its own. A syscall table is an array, indexed by the number in the register, that lists the handler function for each service. And the processor switches to a separate stack, the scratch memory that code uses for its local variables and return addresses, so that kernel code never runs on memory your program controls.
Watch the whole round trip. The seven frames below are one write().
write(fd, buf, 9). So far it's an ordinary function call, and buf holds the 9 bytes of buy milk.Look at which frames depend on the log file. Only the fifth and sixth do, the copy and the store, and for /dev/null the handler does neither: it checks the descriptor and reports all the bytes as written. The registers, the trap, the table lookup and the return happen the same way for any file. So when you write to /dev/null, almost the whole price is the crossing itself.
3.2What 350 nanoseconds adds up to
A write() to /dev/null, which does no real work at all, takes about 348 ns on a recent laptop running macOS. The number depends on the processor, the operating system and the security fixes described in 3.3, so treat it as the right order of magnitude (a few hundred nanoseconds) and measure your own.
?Is that a lot?
Try it on a real case. A service writes each log line with its own write() and logs four million lines a minute. That's four million crossings at 348 ns each, about 1.4 seconds of every minute spent crossing before any bytes go anywhere. If it collects a hundred lines per call instead, there are 40,000 crossings and the cost falls to about 14 milliseconds.
The same arithmetic explains why a syscall per item caps a loop at around three million items a second (one second divided by 350 ns), before any of your own work. The classic ways to hit it are logging with an unbuffered write per line and reading a file a byte at a time. Section 5 is about avoiding both.
3.3Why the price has moved
Old measurements of syscall cost are hard to reproduce, and the reason is a security flaw found in 2018.
Modern processors start working on instructions before they know they're needed, guessing which way the program will branch and throwing the work away if the guess was wrong. This is called speculative execution. In 2018 researchers showed that a user program could use the traces of that thrown-away work to read kernel memory it was never allowed to see. The two families of attack were named Meltdown and Spectre.
The kernel's defence against Meltdown is kernel page table isolation, or KPTI. A page table is the map the processor uses to decide which memory addresses a program may use (chapter 04 covers it). With KPTI, while your program runs in user mode, almost none of the kernel is on that map at all, so there's nothing to read. The price is that every crossing now has to swap the map on the way in and swap it back on the way out. The defences against Spectre add more work at the border, such as clearing the processor's branch-prediction state when control changes sides.
?Why can't you reproduce a syscall benchmark from before 2018?
Because on processors affected by Meltdown, every syscall now does that extra work, and on some workloads the cost of a syscall roughly doubled. Not every processor is affected, so the size of the change depends on the hardware. If an old benchmark's numbers look impossibly good, this is usually the reason and not anything you did wrong. Chapter 44 covers the mitigations themselves.
So a crossing is expensive, and it got dearer. The next question is whether every call that talks to the kernel pays this price in full.
04Calls that skip the door
4.1Data without authority
Some questions a program asks need the kernel's data but none of its authority. The best example is the time. Our logging program asks for it constantly, to stamp every line it writes, and reading a clock changes nothing and needs no permission. Sending every such question through the door would be a waste.
So the kernel keeps a page of memory (the 4 KB or 16 KB unit that page tables hand out) that holds the current time and some code to read it. Each program has its own range of memory addresses, its address space, and the kernel uses the page tables from section 3.3 to place that one page in every program's address space, marked read-only. Your program can read the page like any of its own memory but can't change it. On Linux the page is called the vDSO (virtual dynamic shared object). macOS has the same idea under the name commpage. When your program calls clock_gettime, libc jumps into that page, reads the time the kernel last stored there, adds how far a counter inside the processor, which ticks at a steady rate, has moved since, and returns, all in user mode. Here is that next to the write() from section 3:
This depends on user code being allowed to read that counter. On some virtual machines the counter the kernel uses can't be read from user mode, and there clock_gettime quietly falls back to a real syscall. That's how a program that reads the clock for every log line can run fine on a laptop and spend a surprising share of its time in the kernel on a cloud server.
4.2Calls that never leave libc
The clock is one way for a call to avoid the door. There's another, and it catches people out. Some functions that look like system calls are answered by libc from something it remembered earlier. On macOS, getpid (which returns your process's ID number) is one: libc asks the kernel once at startup and returns the stored number ever after. Linux's glibc, its standard C library, did the same until version 2.25, in 2017, and now does a real syscall.
Before the numbers, a guess.
You time getpid() in a loop on a Mac to measure syscall overhead, and get 1.9 ns per call. What did you measure?
Here are five operations, averaged over a million calls each:
| Call | Cost | Actually a syscall? |
|---|---|---|
| A plain arithmetic loop | 0.22 ns | No: the baseline |
| getpid() | 1.9 ns | **No.** Cached by libc. |
| sbrk(0) | 1.4 ns | **No.** Library-only query. |
| clock_gettime (via steady_clock) | 13.4 ns | No: commpage, kernel data without a trap |
| write(/dev/null, 1 byte) | 347.7 ns | **Yes.** The real thing. |
Only the last row crosses the border. getpid is about 183 times cheaper than write because libc remembers the answer. sbrk(0) asks where the program's memory currently ends and just reads a variable. clock_gettime reads the shared page (the table's timing went through C++'s steady_clock, which calls it). Putting the three ways of getting information out of the kernel side by side:
| Mechanism | Crosses the border? | Typical cost | Example |
|---|---|---|---|
| Real syscall | Yes: trap, check, return | ~350 ns | write, read, openat |
| vDSO or commpage | No: kernel data mapped read-only | ~13 ns | clock_gettime, gettimeofday |
| Cache inside libc | No: libc remembered it | ~2 ns | getpid on macOS |
You can see the same ordering yourself. This C program times three calls the same way, a million each, using clock_gettime as its stopwatch. Save it as cross.c, compile it with clang -O2 cross.c -o cross (use cc on Linux) and run ./cross.
#include <fcntl.h>
#include <stdio.h>
#include <time.h>
#include <unistd.h>
static double ns_per_call(void (*f)(void), long n) {
struct timespec a, b;
clock_gettime(CLOCK_MONOTONIC, &a);
for (long i = 0; i < n; i++) f();
clock_gettime(CLOCK_MONOTONIC, &b);
return ((b.tv_sec - a.tv_sec) * 1e9 + (b.tv_nsec - a.tv_nsec)) / n;
}
static int fd;
static void do_getpid(void) { volatile int p = getpid(); (void)p; }
static void do_clock(void) { struct timespec t; clock_gettime(CLOCK_MONOTONIC, &t); __asm__ volatile("" :: "r"(t.tv_nsec)); }
static void do_write(void) { write(fd, "x", 1); }
int main(void) {
fd = open("/dev/null", O_WRONLY);
long n = 1000000;
printf("getpid() %6.1f ns\n", ns_per_call(do_getpid, n));
printf("clock_gettime() %6.1f ns\n", ns_per_call(do_clock, n));
printf("write(/dev/null, 1 byte) %6.1f ns\n", ns_per_call(do_write, n));
return 0;
}getpid() 1.0 ns
clock_gettime() 16.9 ns
write(/dev/null, 1 byte) 346.2 nsThe absolute numbers move with the machine and with whatever else it's doing, and they land close to the table without matching it. What stays fixed is the ordering. getpid is faster than any trap could be, so it never left the process. clock_gettime costs a few dozen instructions' worth of time, and write costs about twenty times that. If you run this on Linux, expect getpid to cost far more than a nanosecond, much closer to write than to the clock, because after the glibc change above it's a real crossing there.
We now know what a crossing costs and which calls avoid it. The ones that can't avoid it will always cost a few hundred nanoseconds each, so the only thing left to change is how many of them we make.
05Making fewer trips
5.1Collecting lines into one write
Take our logging service at 500,000 lines a second, one write() per line. That's 500,000 crossings at 348 ns each, which is 174 ms of every second. A processor is made of several cores, each able to run one thread at a time, so this means 17.4% of one core's time goes on crossing, before any line has reached a disk.
The fix follows from the cost. If a crossing costs the same whether it carries one line or forty, carry forty. The program keeps a buffer, a block of its own memory, and copies each line into that instead of calling the kernel. Only when the buffer is full does it make one write() for everything in it. Copying into your own memory involves no crossing, so each line costs a few nanoseconds. With log lines of about 100 bytes, a 4 KB buffer holds about forty of them.
write() for it now would cost a whole crossing.Almost every logging library does this, and so does stdio, the C library's buffered input and output functions, including printf. A printf in a loop that writes to a file or a pipe isn't a syscall per call, because stdio collects the output and flushes it in blocks. When the output goes to a terminal, stdio flushes at every newline instead, so that you see each line as it's printed.
Put together, the whole calculation looks like this:
| Log lines per second | a busy service | 500,000 |
| Unbuffered: one write each | 500,000 × 348 ns | 174 ms/s |
| CPU spent crossing | 174 ms of every second | 17.4% |
| Buffered to 4 KB, ~40 lines per write | 12,500 × 348 ns | 4.3 ms/s |
| CPU recovered by adding a buffer | ≈ 17% of one core | |
5.2Other ways to do more per crossing
Buffering collects many small writes into one. The same idea works in other shapes, and each of the following removes crossings in a different way.
| Technique | What it does | Crossings |
|---|---|---|
| Buffered I/O | stdio accumulates writes and flushes in blocks | One per block |
| Vectored I/O | writev and readv take an array of buffers, so scattered pieces go out in one call | One per array |
sendfile and splice | Move bytes between two descriptors inside the kernel, so they never come up into your program and back down | One per transfer, no copy into your program |
| Batched calls | sendmmsg and recvmmsg send or receive many network messages at once, and epoll lets a server ask about thousands of connections in one call, which is how the nginx web server and the Redis database cope | One per batch |
| io_uring | Two queues in memory shared with the kernel, one for requests and one for results | One per batch, or potentially zero in the polled mode of 5.3 |
The first four are small changes to how you call the kernel. The last one changes the shape of the whole conversation, and it's worth looking at on its own.
5.3io_uring: a queue instead of a door
Buffering and batching make fewer trips. Jens Axboe's io_uring, added to Linux in 2019, goes further and tries to make trips unnecessary. It sets up two ring buffers in memory that both your program and the kernel can see. A ring buffer is a fixed-size queue where entries are added at one end and taken from the other, wrapping around when it reaches the end of its memory. One ring carries your requests to the kernel (submission queue entries, SQEs) and the other carries results back (completion queue entries, CQEs). Both sides track their position with a pointer, a head for where the next entry is taken and a tail for where the next one is added.
In the ordinary mode, your program fills in as many requests as it likes and then makes one syscall, io_uring_enter, to tell the kernel they're there, so a whole batch costs one crossing. With one setup option, IORING_SETUP_SQPOLL, even that call goes away, because a kernel thread watches the ring for new entries:
?So why isn't everything built on io_uring?
Because you pay for it elsewhere: a much harder API, a newer kernel (Linux 5.10 or later is a sensible floor), and, in polled mode, a kernel thread keeping a core busy checking the ring. It changes the shape of your code, so it belongs last in the list of fixes in section 8.2.
Everything in this section has been the program deciding when to cross. The door also opens from the other side, when the kernel comes in without being asked.
06Interrupts: the door from the other side
6.1A packet arrives while your code runs
Suppose that while your program is in the middle of formatting a log line, a network packet arrives for some other program on the machine. A packet is one small chunk of data sent over a network, and it arrives at the NIC, the network interface card, the hardware that sends and receives packets. Your code has no idea any of this is happening, and the NIC can't wait for your code to ask.
So the hardware forces the issue. The NIC raises an interrupt, a notice from the hardware that makes the processor stop what it's doing and enter the kernel, exactly as a trap would, only without your program asking. A timer chip does the same thing at regular intervals, which is how the kernel gets to switch between programs. Here's what happens to your thread:
?Why split an interrupt handler in two?
Because interrupts are disabled or restricted while a handler runs, so the handler has to be short, or the machine can't respond to anything else for as long as it takes. Linux splits the work. The top half acknowledges the device and queues work, and the bottom half (which Linux implements as softirqs, tasklets or threaded handlers) does the real processing later, with interrupts enabled.
6.2NAPI: when interrupts are the problem
Under heavy load there's a flaw in this design. If packets arrive faster than the machine can process them, it spends all its time entering and leaving interrupt handlers, and user programs, yours included, get no processor time at all.
Linux's answer for network cards is called NAPI. When the first packet's interrupt arrives, the driver turns the NIC's interrupts off and switches to polling: the kernel keeps asking the NIC "is there anything yet?" and handles whatever has arrived in batches. Only when the NIC has nothing left does the driver turn interrupts back on. A quiet machine still gets one interrupt per packet, and a flooded one gets very few. That trades latency, how long each packet waits before it's handled, for the ability to make progress at all. Chapter 10 follows the poll loop through the network stack.
An interrupt stops your thread so the kernel can run. There's a mechanism that does the reverse, stopping your thread so that some of your code can run, and it carries a rule that's easy to break.
07Signals, and what a handler may do
7.1Interrupting your own code
Suppose someone presses Ctrl-C while your logging program is running. The kernel tells your program about it with a signal, a numbered notice of an event; Ctrl-C sends the one called SIGINT. By default that signal ends the program, but you can register a function to run instead, a signal handler. The kernel doesn't wait for a convenient moment to run it. It stops the thread wherever it happens to be, runs the handler, and then lets the thread carry on. The handler runs on the stack of the thread it interrupted (unless you set up a separate one with sigaltstack), at whatever instruction that thread was on.
That last fact is the trouble. Suppose your handler wants to log a final line, interrupted, buy milk unsaved, using printf. Look at what that can do:
malloc, which hands out memory from the heap, the region your program draws memory from. It takes the allocator's lock, a flag only one thread can hold at a time so that nobody else changes the heap's bookkeeping meanwhile, and starts updating its free lists.?Why is the list of safe calls inside a handler so short?
Because the handler may have interrupted the very function it wants to call, midway through updating that function's state. So POSIX, the standard that Unix-like systems follow, publishes a short list of async-signal-safe functions, ones written so that calling them at any moment is harmless. It excludes malloc, printf and almost everything else. It does include write, the very call this chapter has been timing, so a handler can log a fixed message with one write() and no buffer.
The usual pattern is to do almost nothing in the handler. It sets a flag, declared volatile sig_atomic_t (a type the standard guarantees can be set in one uninterruptible step), or writes one byte into a pipe, a pair of descriptors where bytes written to one end can be read from the other. Then the main loop notices the flag or the byte and does the real work, which is a safe place to call printf. A pipe the program uses to talk to itself this way is called a self-pipe.
That's every way across the border, in both directions. What's left is finding out how often a real program crosses, and what to do about it.
08Counting your crossings
8.1Finding the crossings
Each question this chapter raised has a tool that answers it on a running program.
# How many crossings, and which calls? (sections 3 and 5) Start here.
strace -c -f ./app # Linux
dtruss -c ./app # macOS (needs System Integrity Protection relaxed)
# Which call site made them? (section 5)
perf trace ./app # Linux
perf record -e raw_syscalls:sys_enter -g ./app
# How much time goes to the kernel rather than your code? (section 3)
/usr/bin/time -v ./app 2>&1 | grep -E 'System time|User time' # Linux; macOS: /usr/bin/time -l
# Is the interrupt load the problem? (section 6, Linux)
cat /proc/interrupts
mpstat -P ALL 1 # %irq and %soft columnsstrace -c is the first command to run, and most of the time it answers the question. A count of 4 million write calls where you expected a few thousand is the entire diagnosis.
8.2Reducing crossings, in order
- Buffer. The biggest win for the least effort, with no new API (section 5.1).
- Batch.
writev,readv,sendmmsgandrecvmmsgmake one crossing carry many items. - Avoid the copy.
sendfile,splice, ormmap(which maps a file into your program's memory) for large file reads, so the data doesn't travel up into your program and back down. - io_uring, on Linux 5.10 or newer, when you need hundreds of thousands of operations per second. It changes the shape of your code, so it belongs last.
8.3What you trade for what
| You get | You pay | When the bill arrives |
|---|---|---|
| A protected kernel that checks everything | ~350 ns per crossing | In any loop that calls per item |
| vDSO and cached values for cheap calls | You can't tell which is which by reading C | When you benchmark getpid and measure nothing |
| Buffered I/O, far fewer crossings | Data sits in your program's memory until flushed | On a crash, as lost log lines |
| io_uring, potentially zero syscalls | A harder API and a newer kernel | As complexity, up front |
| Spectre and Meltdown mitigations keeping you safe | Every crossing got more expensive from 2018 on affected processors | As old benchmarks that won't reproduce |
8.4Symptom, cause, fix
| Symptom | Likely cause | Fix |
|---|---|---|
| High system time you didn't expect | A syscall per item: unbuffered logging, byte-at-a-time reads | strace -c -f to find it, then buffer |
| A syscall benchmark reads ~2 ns | The call is cached in libc and never crosses (getpid on macOS) | Benchmark write to /dev/null instead |
| An old benchmark won't reproduce | KPTI and Spectre mitigations since 2018 | Compare against a post-2018 baseline |
High %irq or %soft in mpstat | Interrupt load starving user threads | Check /proc/interrupts; see NAPI and chapter 10 |
| Process hangs inside a signal handler | A call that isn't async-signal-safe, like malloc or printf | Set a flag or write to a self-pipe |
09Summary
- User code can't touch hardware. Those instructions fault in user mode, so every disk and network request goes through the kernel, as
buy milkdid in section 3. - There's one door in. A trap instruction jumps to a single entry point, and a number in a register picks the service.
- The kernel trusts nothing you pass. It checks every pointer and copies the data across, which is where much of the cost comes from.
- A real crossing costs about 350 ns, even for a
writeto/dev/nullthat does no work. A million of them is a third of a second. - Some calls never cross.
clock_gettimereads the vDSO or commpage in about 13 ns, andgetpidon macOS returns a cached value in about 2 ns. - You can't tell which is which from the C. Check with
strace -con Linux ordtruss -con macOS. - Mitigations made every crossing dearer in 2018. On affected processors, KPTI roughly doubled syscall cost on some workloads.
- Fast code crosses less. Buffering forty log lines into one
writecut 17.4% of a core to about 0.4%. After buffering come batching, avoiding the copy, and io_uring, in that order. - Interrupts use the same door from the other side. A short top half acknowledges the device and the bottom half does the work with interrupts enabled.
- Signal handlers may call almost nothing. Set a flag or write to a self-pipe, and do the work in the main loop.
10Build this
Measure your own border, and find a call that isn't crossing it. You already ran the first half of this in sections 1 and 4.
- Time
getpid,clock_gettime,writeto/dev/null, and an empty loop. - Confirm with
strace -c(Linux) ordtruss -c(macOS) which ones produce a real syscall. At least one will surprise you. - Then write one megabyte to a file, one byte at a time, and again in 4 KB blocks. Compare wall time and the
strace -ccounts.
Your 4 KB version should make about four thousand times fewer crossings (a million against 245) and run much faster. Having produced that ratio yourself makes buffering an instinct instead of advice.
11Interview questions
beginnerRoughly what does a syscall cost, and why isn't it free?›
A few hundred nanoseconds: about 350 for a write to /dev/null on a recent laptop, with the exact figure depending on the processor, the operating system and the security mitigations. The cost comes from the controlled crossing: switch privilege level, change stacks, dispatch through a table, check every pointer the caller passed, do the work, then restore everything. Checking matters because user code could change a pointer concurrently, so the kernel copies instead of trusting addresses.
intermediateWhy is calling clock_gettime so much cheaper than calling write?›
It usually isn't a syscall. Linux and macOS both map a small read-only page into every process (the vDSO on Linux, the commpage on macOS) containing code and a timestamp the kernel keeps current. Calling clock_gettime jumps into that page and reads the value without ever trapping.
That's about 13 ns against about 350 for write, roughly twenty-six times cheaper. The C looks the same either way, but one call crosses the border and the other doesn't.
intermediateYour service spends 20% of CPU in the kernel and you expected almost none. First step?›
strace -c -f. It gives you syscall counts and time by type in one command, and the answer is usually obvious: a few million write calls from unbuffered logging, or a clock read per loop iteration on a virtual machine where the clock can't be read from the shared page.
At roughly 350 ns each, half a million log lines a second is about 17% of a core spent purely crossing. Buffering to 4 KB blocks cuts that by a factor of forty.
deepWhat may you safely call inside a signal handler, and why is the list so short?›
Only async-signal-safe functions, a short POSIX list. A handler runs at an arbitrary instruction on the interrupted thread's stack, possibly while the interrupted code holds a lock or is midway through updating allocator state.
Call malloc from a handler that interrupted malloc and you corrupt the heap or deadlock. The safe pattern is to set a volatile sig_atomic_t flag or write one byte to a self-pipe, then do the real work from the main loop.
deepHow does io_uring reduce syscall cost to nearly zero?›
By replacing per-operation crossings with shared memory. Two ring buffers, submission and completion, are mapped into both the process and the kernel. The application writes submission entries directly into memory the kernel can already see and advances a tail pointer.
With IORING_SETUP_SQPOLL a kernel thread polls that ring, so the data path can run with no syscalls at all. You trade a much harder API and a dedicated polling thread for removing the border from the hot path entirely.
12Go deeper
You benchmark syscall overhead with getpid and measure 2 ns. What went wrong?›
On macOS, and on Linux with older glibc, getpid isn't a syscall: libc caches the value at process start. Use something that must reach the kernel, like write to /dev/null, and confirm with strace -c or dtruss -c.
Why can't you call printf in a signal handler?›
It isn't async-signal-safe. The handler may have interrupted stdio or the allocator mid-update, so calling back in can deadlock or corrupt state.
Unbuffered logging at 500k lines/sec. Roughly what does it cost?›
About 17% of a core in crossings alone, at ~350 ns each. Buffering to 4 KB cuts it to under half a percent.
Why did syscalls get slower in 2018?›
Spectre and Meltdown mitigations: page table isolation and predictor flushing across the border. Old benchmarks stopped reproducing.
"Mechanism: Limited Direct Execution" builds the user mode, kernel mode and trap instruction from this chapter one step at a time. Free online at ostep.org.
Design document from the author. Short, and the clearest explanation of the ring protocol.
What's in the mapped page and why. Explains the 13 ns row in section 4.
Actual list of what you may call in a handler. Shorter than you expect.
Lower-overhead than strace and usable on production traffic, which strace mostly is not.
13Related chapters
Interrupts, softirqs and NAPI at full depth, following one packet up to
your read(). Chapter 10.
Why crossings got more expensive in 2018, and what KPTI does. Chapter 44.
What happens after read() and write() cross the border.
Chapter 08.
Tracing syscalls on a live system with less overhead than strace. Chapter 48.