KnowSys

Syscalls, Interrupts & the Kernel Boundary

Follow one line of output, `buy milk`, through the narrow door between your program and the kernel: why the call costs a few hundred nanoseconds even when it does no work, which calls never use the door, and how fast programs make fewer trips.

⏱ 27 min read◆ BeginnerAssumes: a terminal; CPU basics and virtual memory help
Start reading

Your program has a line to record, buy milk, and it calls write() to add that line to a log file. It's one line of code, and it returns almost at once. Run it again for the next request, and again for the one after, and a busy service ends up doing it hundreds of thousands of times a second.

Now look at what that line asks for. A log file lives on a disk, and your program can't reach the disk. The processor won't let ordinary program code touch disks, network cards, or the memory of other programs. The code that is allowed to is the kernel, the core of the operating system, which runs with powers your program doesn't have. So when your program calls write(), someone else does the work: the request travels to the kernel, the kernel decides whether to honour it, does the writing, and reports back.

Each of those trips is quick, and a million of them are not. This chapter follows that one write() of buy milk through the trip and asks three things: what happens on the way, what does it cost, and how do fast programs avoid paying it so often? We start by timing it, because the answer to the second question is a surprise.

01Timing a million small writes

1.1A million tiny trips against a few big ones

Here's the smallest experiment that shows the whole chapter. It sends about a megabyte to a place that throws everything away, once as a million separate one-byte calls, and once as a few hundred calls of 4 KB each (4,096 bytes).

A few pieces of the script need introducing. When a program opens a file, the kernel hands back a file descriptor, a small number the program passes to every later call to say which file it means. os.open returns one here, and os.write(fd, data) is Python's thin wrapper around the kernel's write call. The file being opened is /dev/null, a special file that accepts any bytes you write to it and discards them. Using it means no disk is involved, so any time the script spends belongs to the calls themselves. time.perf_counter() reads a stopwatch in seconds, and the script reads it before and after each loop.

Save it as cross.py and run it with python3 cross.py. Any Mac or Linux machine will do.

Write the same 1,000,000 bytes to /dev/null two ways
python
Python
import os, time
fd = os.open("/dev/null", os.O_WRONLY)    # a file that throws away what you write
n = 1_000_000
 
t = time.perf_counter()
for _ in range(n):
    os.write(fd, b"x")                    # one byte per call
a = time.perf_counter() - t
 
block = b"x" * 4096
t = time.perf_counter()
for _ in range(n // 4096 + 1):
    os.write(fd, block)                   # 4 KB per call
b = time.perf_counter() - t
 
print(f"1,000,000 writes of 1 byte : {a*1000:7.1f} ms")
print(f"      245 writes of 4 KB   : {b*1000:7.1f} ms")
output
C++
1,000,000 writes of 1 byte :   378.1 ms
      245 writes of 4 KB   :     0.1 ms

Both versions send the same megabyte. The first makes a million calls and takes over a third of a second. The second makes 245 calls and finishes in about a tenth of a millisecond, thousands of times sooner. The exact times change from machine to machine and from run to run, but the second line always comes out thousands of times smaller than the first.

1.2What the numbers say

Nothing was stored anywhere, since /dev/null discards its input, so the time went into the calls themselves. Divide 378 milliseconds by a million and each call cost about 380 nanoseconds, where a nanosecond is a billionth of a second. That figure includes Python's own loop, so the trip itself is a little cheaper: section 3 puts it at about 350.

Three hundred and fifty nanoseconds is nothing once. A million times, it's a third of a second, and fast programs are the ones that make fewer trips with more in each. But a call that does no work at all still spends hundreds of nanoseconds, and that needs explaining. To see where the time goes, we first have to ask why there's a border to cross.

02Why there's a border at all

2.1Two privilege modes

?Why can't your program just talk to the disk itself?

Because then every program could. A bug in your text editor could scribble over the filesystem, and any program could read another program's memory or reprogram the network card to watch other people's traffic. Something has to stand between programs and the hardware.

The processor enforces this itself. It runs in one of (at least) two modes. In user mode, where your code runs, the instructions that talk to hardware or change which memory belongs to which program are forbidden. If your program tries one, the processor stops it with a fault and hands control to the kernel, which usually ends the program. In kernel mode those instructions work, and only the kernel runs there. The mode a piece of code runs in is its privilege level. (ARM processors, such as the ones in recent Macs, call these levels EL0 for user code and EL1 for the kernel.)

Four nested circles labelled Ring 0 (kernel) at the centre, Rings 1 and 2 (device drivers) and Ring 3 (applications) outermost, with a scale from most to least privileged
x86 chips actually define four privilege levels, drawn as rings with the kernel at the centre. Linux, Windows and macOS use only two: ring 0 is kernel mode and ring 3 is user mode. The two middle rings, intended for device drivers, go unused, and drivers on these systems run in ring 0 with the rest of the kernel.Image: Hertzsprung at English Wikipedia, CC BY-SA 3.0, via Wikimedia Commons

2.2The one door in

That leaves a puzzle. If your program can't touch the disk, how does it ever read a file?

Think of a bank vault with a teller window. You can't walk into the vault, and the bank doesn't want other customers walking in either. So you fill out a slip, hand it through the window, and wait while the teller checks it, fetches what you asked for, and passes it back. Your program asks the kernel for things the same way, and the processor provides the window.

?Why not let programs call into the kernel like any other function?

Because a program that could jump to any address in the kernel could jump past the line that checks its permissions. So the processor offers exactly one way in: a special instruction (syscall on x86-64 chips, svc on ARM chips such as Apple's), which we'll call the trap instruction. Before running it, your program puts a number in a register, one of the handful of tiny storage slots inside the processor, to say which service it wants: read, write, or openat (open a file). The arguments go in other registers. The trap instruction switches the processor to kernel mode and jumps to a single entry point that the kernel chose when the machine booted.

That whole request through the door is a system call, or syscall. It's a controlled transfer into code you can't see, running at a privilege level you don't have. The kernel promises to check every argument, do the thing, and hand control back, or to fail with an error code and leave the system intact.

A large grid map of the Linux kernel: columns for human interface, system, processing, memory, storage and networking, rows from user space interfaces at the top down to hardware at the bottom
What sits behind the door. Each column of this map of the Linux kernel is one area of its work: processes, memory, storage, networking and so on. The top row is the system calls, the only way in from user space. Everything below it, down to the hardware in the bottom row, runs in kernel mode.Image: Constantine Shulyupin, CC BY-SA 4.0, via Wikimedia Commons

2.3Why the kernel copies your arguments

The kernel can't trust anything that comes through the door, and the arguments are the first place that shows.

?Why doesn't the kernel just use the pointer you gave it?

A pointer is an address in memory where data lives, and write takes one to say where your bytes are. A pointer you hand over might point into the kernel's own memory, which would trick the kernel into reading its secrets for you. And your program may have several threads, separate flows of execution that share memory, so another thread could change the buffer while the kernel is halfway through reading it.

So the kernel checks every pointer and length, then copies the bytes into memory of its own and works from the copy. That copy is a real cost, which is part of why large transfers behave differently from small ones.

That's the design. Now we can follow our one write() of buy milk through it, step by step, and put a price on the trip.

03Inside one write()

3.1One write(), across the border and back

Our program calls write(fd, buf, 9). The first argument is the descriptor of the log file; say the program opened the log earlier and the kernel handed back descriptor 3. buf is a pointer to the line itself, the eight characters of buy milk followed by the newline that ends the line, and 9 is how many bytes that makes. The function your code calls isn't the kernel's. It lives in libc, the standard C library that nearly every program on a Unix-like system links in, and it's a small wrapper whose only job is to set up the trap.

Two applications on the right; arrows labelled function calls go into a green ring labelled GNU C Library, and arrows labelled system calls go into a red inner ring labelled System Call Interface, around the kernel's memory manager, scheduler, file system, network and I/O parts
The same layering on Linux. Programs mostly make function calls into glibc, Linux's usual C library, and glibc's wrappers make the system calls. A program can also make a system call directly, without libc. Either way, the red System Call Interface ring is the only way to reach the kernel's parts in the middle. The drawing is from 2013, and Linux has added system calls since.Image: ScotXW, CC BY-SA 3.0, via Wikimedia Commons

The kernel also needs two things of its own. A syscall table is an array, indexed by the number in the register, that lists the handler function for each service. And the processor switches to a separate stack, the scratch memory that code uses for its local variables and return addresses, so that kernel code never runs on memory your program controls.

Watch the whole round trip. The seven frames below are one write().

One write(fd, buf, 9) across the kernel boundary
Your programuser mode · can't touch hardwareRegistersthe number and the argumentsKernelkernel mode · full powersLog filekept by the kernelbufbuy milk\nwritesyscall numberfd 3which file&bufwhere the bytes arelen 9how manyentry pointchosen at bootkernel copybuy milk\nbuy milkresult 9bytes written
Step 1. Your code calls write(fd, buf, 9). So far it's an ordinary function call, and buf holds the 9 bytes of buy milk.
1 / 7

Look at which frames depend on the log file. Only the fifth and sixth do, the copy and the store, and for /dev/null the handler does neither: it checks the descriptor and reports all the bytes as written. The registers, the trap, the table lookup and the return happen the same way for any file. So when you write to /dev/null, almost the whole price is the crossing itself.

3.2What 350 nanoseconds adds up to

A write() to /dev/null, which does no real work at all, takes about 348 ns on a recent laptop running macOS. The number depends on the processor, the operating system and the security fixes described in 3.3, so treat it as the right order of magnitude (a few hundred nanoseconds) and measure your own.

?Is that a lot?

Try it on a real case. A service writes each log line with its own write() and logs four million lines a minute. That's four million crossings at 348 ns each, about 1.4 seconds of every minute spent crossing before any bytes go anywhere. If it collects a hundred lines per call instead, there are 40,000 crossings and the cost falls to about 14 milliseconds.

The same arithmetic explains why a syscall per item caps a loop at around three million items a second (one second divided by 350 ns), before any of your own work. The classic ways to hit it are logging with an unbuffered write per line and reading a file a byte at a time. Section 5 is about avoiding both.

3.3Why the price has moved

Old measurements of syscall cost are hard to reproduce, and the reason is a security flaw found in 2018.

Modern processors start working on instructions before they know they're needed, guessing which way the program will branch and throwing the work away if the guess was wrong. This is called speculative execution. In 2018 researchers showed that a user program could use the traces of that thrown-away work to read kernel memory it was never allowed to see. The two families of attack were named Meltdown and Spectre.

The kernel's defence against Meltdown is kernel page table isolation, or KPTI. A page table is the map the processor uses to decide which memory addresses a program may use (chapter 04 covers it). With KPTI, while your program runs in user mode, almost none of the kernel is on that map at all, so there's nothing to read. The price is that every crossing now has to swap the map on the way in and swap it back on the way out. The defences against Spectre add more work at the border, such as clearing the processor's branch-prediction state when control changes sides.

?Why can't you reproduce a syscall benchmark from before 2018?

Because on processors affected by Meltdown, every syscall now does that extra work, and on some workloads the cost of a syscall roughly doubled. Not every processor is affected, so the size of the change depends on the hardware. If an old benchmark's numbers look impossibly good, this is usually the reason and not anything you did wrong. Chapter 44 covers the mitigations themselves.

So a crossing is expensive, and it got dearer. The next question is whether every call that talks to the kernel pays this price in full.

04Calls that skip the door

4.1Data without authority

Some questions a program asks need the kernel's data but none of its authority. The best example is the time. Our logging program asks for it constantly, to stamp every line it writes, and reading a clock changes nothing and needs no permission. Sending every such question through the door would be a waste.

So the kernel keeps a page of memory (the 4 KB or 16 KB unit that page tables hand out) that holds the current time and some code to read it. Each program has its own range of memory addresses, its address space, and the kernel uses the page tables from section 3.3 to place that one page in every program's address space, marked read-only. Your program can read the page like any of its own memory but can't change it. On Linux the page is called the vDSO (virtual dynamic shared object). macOS has the same idea under the name commpage. When your program calls clock_gettime, libc jumps into that page, reads the time the kernel last stored there, adds how far a counter inside the processor, which ticks at a steady rate, has moved since, and returns, all in user mode. Here is that next to the write() from section 3:

Reading the clock without a trap, then a write() that needs one
Your programuser modeShared pageread-only · mapped into your programKernelkernel mode · ~350 ns to enterkernel clockkeeps tickingtime10:00:00.001now10:00:00.002write handlerthe whole trip
Step 1. The kernel keeps its clock and a copy of the time in a page that every program can read but not change.
1 / 6

This depends on user code being allowed to read that counter. On some virtual machines the counter the kernel uses can't be read from user mode, and there clock_gettime quietly falls back to a real syscall. That's how a program that reads the clock for every log line can run fine on a laptop and spend a surprising share of its time in the kernel on a cloud server.

4.2Calls that never leave libc

The clock is one way for a call to avoid the door. There's another, and it catches people out. Some functions that look like system calls are answered by libc from something it remembered earlier. On macOS, getpid (which returns your process's ID number) is one: libc asks the kernel once at startup and returns the stored number ever after. Linux's glibc, its standard C library, did the same until version 2.25, in 2017, and now does a real syscall.

Before the numbers, a guess.

Predict before you read on

You time getpid() in a loop on a Mac to measure syscall overhead, and get 1.9 ns per call. What did you measure?

Here are five operations, averaged over a million calls each:

CallCostActually a syscall?
A plain arithmetic loop0.22 nsNo: the baseline
getpid()1.9 ns**No.** Cached by libc.
sbrk(0)1.4 ns**No.** Library-only query.
clock_gettime (via steady_clock)13.4 nsNo: commpage, kernel data without a trap
write(/dev/null, 1 byte)347.7 ns**Yes.** The real thing.

Only the last row crosses the border. getpid is about 183 times cheaper than write because libc remembers the answer. sbrk(0) asks where the program's memory currently ends and just reads a variable. clock_gettime reads the shared page (the table's timing went through C++'s steady_clock, which calls it). Putting the three ways of getting information out of the kernel side by side:

MechanismCrosses the border?Typical costExample
Real syscallYes: trap, check, return~350 nswrite, read, openat
vDSO or commpageNo: kernel data mapped read-only~13 nsclock_gettime, gettimeofday
Cache inside libcNo: libc remembered it~2 nsgetpid on macOS

You can see the same ordering yourself. This C program times three calls the same way, a million each, using clock_gettime as its stopwatch. Save it as cross.c, compile it with clang -O2 cross.c -o cross (use cc on Linux) and run ./cross.

Time getpid, clock_gettime and write to /dev/null
c
C
#include <fcntl.h>
#include <stdio.h>
#include <time.h>
#include <unistd.h>
 
static double ns_per_call(void (*f)(void), long n) {
    struct timespec a, b;
    clock_gettime(CLOCK_MONOTONIC, &a);
    for (long i = 0; i < n; i++) f();
    clock_gettime(CLOCK_MONOTONIC, &b);
    return ((b.tv_sec - a.tv_sec) * 1e9 + (b.tv_nsec - a.tv_nsec)) / n;
}
 
static int fd;
static void do_getpid(void) { volatile int p = getpid(); (void)p; }
static void do_clock(void)  { struct timespec t; clock_gettime(CLOCK_MONOTONIC, &t); __asm__ volatile("" :: "r"(t.tv_nsec)); }
static void do_write(void)  { write(fd, "x", 1); }
 
int main(void) {
    fd = open("/dev/null", O_WRONLY);
    long n = 1000000;
    printf("getpid()                 %6.1f ns\n", ns_per_call(do_getpid, n));
    printf("clock_gettime()          %6.1f ns\n", ns_per_call(do_clock, n));
    printf("write(/dev/null, 1 byte) %6.1f ns\n", ns_per_call(do_write, n));
    return 0;
}
output
C++
getpid()                    1.0 ns
clock_gettime()            16.9 ns
write(/dev/null, 1 byte)  346.2 ns

The absolute numbers move with the machine and with whatever else it's doing, and they land close to the table without matching it. What stays fixed is the ordering. getpid is faster than any trap could be, so it never left the process. clock_gettime costs a few dozen instructions' worth of time, and write costs about twenty times that. If you run this on Linux, expect getpid to cost far more than a nanosecond, much closer to write than to the clock, because after the glibc change above it's a real crossing there.

We now know what a crossing costs and which calls avoid it. The ones that can't avoid it will always cost a few hundred nanoseconds each, so the only thing left to change is how many of them we make.

05Making fewer trips

5.1Collecting lines into one write

Take our logging service at 500,000 lines a second, one write() per line. That's 500,000 crossings at 348 ns each, which is 174 ms of every second. A processor is made of several cores, each able to run one thread at a time, so this means 17.4% of one core's time goes on crossing, before any line has reached a disk.

The fix follows from the cost. If a crossing costs the same whether it carries one line or forty, carry forty. The program keeps a buffer, a block of its own memory, and copies each line into that instead of calling the kernel. Only when the buffer is full does it make one write() for everything in it. Copying into your own memory involves no crossing, so each line costs a few nanoseconds. With log lines of about 100 bytes, a 4 KB buffer holds about forty of them.

Forty log lines, one crossing
Your programmakes a lineBufferyour own memory · no crossing · 4 KBKernel~350 ns per crossingline 1buy milkline 2buy eggsline 3buy tea37 morelines40 lines4 KB, one tripline 41not flushed yet
Step 1. Your program produces its first log line. Calling write() for it now would cost a whole crossing.
1 / 6

Almost every logging library does this, and so does stdio, the C library's buffered input and output functions, including printf. A printf in a loop that writes to a file or a pipe isn't a syscall per call, because stdio collects the output and flushes it in blocks. When the output goes to a terminal, stdio flushes at every newline instead, so that you see each line as it's printed.

Put together, the whole calculation looks like this:

Log lines per seconda busy service500,000
Unbuffered: one write each500,000 × 348 ns174 ms/s
CPU spent crossing174 ms of every second17.4%
Buffered to 4 KB, ~40 lines per write12,500 × 348 ns4.3 ms/s
CPU recovered by adding a buffer≈ 17% of one core

5.2Other ways to do more per crossing

Buffering collects many small writes into one. The same idea works in other shapes, and each of the following removes crossings in a different way.

TechniqueWhat it doesCrossings
Buffered I/Ostdio accumulates writes and flushes in blocksOne per block
Vectored I/Owritev and readv take an array of buffers, so scattered pieces go out in one callOne per array
sendfile and spliceMove bytes between two descriptors inside the kernel, so they never come up into your program and back downOne per transfer, no copy into your program
Batched callssendmmsg and recvmmsg send or receive many network messages at once, and epoll lets a server ask about thousands of connections in one call, which is how the nginx web server and the Redis database copeOne per batch
io_uringTwo queues in memory shared with the kernel, one for requests and one for resultsOne per batch, or potentially zero in the polled mode of 5.3

The first four are small changes to how you call the kernel. The last one changes the shape of the whole conversation, and it's worth looking at on its own.

5.3io_uring: a queue instead of a door

Buffering and batching make fewer trips. Jens Axboe's io_uring, added to Linux in 2019, goes further and tries to make trips unnecessary. It sets up two ring buffers in memory that both your program and the kernel can see. A ring buffer is a fixed-size queue where entries are added at one end and taken from the other, wrapping around when it reaches the end of its memory. One ring carries your requests to the kernel (submission queue entries, SQEs) and the other carries results back (completion queue entries, CQEs). Both sides track their position with a pointer, a head for where the next entry is taken and a tail for where the next one is added.

In the ordinary mode, your program fills in as many requests as it likes and then makes one syscall, io_uring_enter, to tell the kernel they're there, so a whole batch costs one crossing. With one setup option, IORING_SETUP_SQPOLL, even that call goes away, because a kernel thread watches the ring for new entries:

An operation through io_uring with SQPOLL
Your appSubmission ringKernel poll threadCompletion ringwrite SQEadvance tailpollpost CQEread CQE
Step 1. The app writes a submission entry directly into memory the kernel can already see. No syscall.
1 / 5

?So why isn't everything built on io_uring?

Because you pay for it elsewhere: a much harder API, a newer kernel (Linux 5.10 or later is a sensible floor), and, in polled mode, a kernel thread keeping a core busy checking the ring. It changes the shape of your code, so it belongs last in the list of fixes in section 8.2.

Everything in this section has been the program deciding when to cross. The door also opens from the other side, when the kernel comes in without being asked.

06Interrupts: the door from the other side

6.1A packet arrives while your code runs

Suppose that while your program is in the middle of formatting a log line, a network packet arrives for some other program on the machine. A packet is one small chunk of data sent over a network, and it arrives at the NIC, the network interface card, the hardware that sends and receives packets. Your code has no idea any of this is happening, and the NIC can't wait for your code to ask.

So the hardware forces the issue. The NIC raises an interrupt, a notice from the hardware that makes the processor stop what it's doing and enter the kernel, exactly as a trap would, only without your program asking. A timer chip does the same thing at regular intervals, which is how the kernel gets to switch between programs. Here's what happens to your thread:

A packet interrupts your thread
NIChardwareCPU corerunning your threadTop halfkernel · must be shortBottom halfkernel · does the real workyour coderunningpackethandlerinterrupts held backwork itemprocess the packet
Step 1. Your thread is running ordinary code, formatting a log line. A packet arrives at the NIC, for some other program.
1 / 6

?Why split an interrupt handler in two?

Because interrupts are disabled or restricted while a handler runs, so the handler has to be short, or the machine can't respond to anything else for as long as it takes. Linux splits the work. The top half acknowledges the device and queues work, and the bottom half (which Linux implements as softirqs, tasklets or threaded handlers) does the real processing later, with interrupts enabled.

6.2NAPI: when interrupts are the problem

Under heavy load there's a flaw in this design. If packets arrive faster than the machine can process them, it spends all its time entering and leaving interrupt handlers, and user programs, yours included, get no processor time at all.

Linux's answer for network cards is called NAPI. When the first packet's interrupt arrives, the driver turns the NIC's interrupts off and switches to polling: the kernel keeps asking the NIC "is there anything yet?" and handles whatever has arrived in batches. Only when the NIC has nothing left does the driver turn interrupts back on. A quiet machine still gets one interrupt per packet, and a flooded one gets very few. That trades latency, how long each packet waits before it's handled, for the ability to make progress at all. Chapter 10 follows the poll loop through the network stack.

An interrupt stops your thread so the kernel can run. There's a mechanism that does the reverse, stopping your thread so that some of your code can run, and it carries a rule that's easy to break.

07Signals, and what a handler may do

7.1Interrupting your own code

Suppose someone presses Ctrl-C while your logging program is running. The kernel tells your program about it with a signal, a numbered notice of an event; Ctrl-C sends the one called SIGINT. By default that signal ends the program, but you can register a function to run instead, a signal handler. The kernel doesn't wait for a convenient moment to run it. It stops the thread wherever it happens to be, runs the handler, and then lets the thread carry on. The handler runs on the stack of the thread it interrupted (unless you set up a separate one with sigaltstack), at whatever instruction that thread was on.

That last fact is the trouble. Suppose your handler wants to log a final line, interrupted, buy milk unsaved, using printf. Look at what that can do:

A signal handler that calls malloc
Main codeAllocator lockKernelSignal handlerlocksignalrun handlerlock?wait forever
Step 1. Your main code calls malloc, which hands out memory from the heap, the region your program draws memory from. It takes the allocator's lock, a flag only one thread can hold at a time so that nobody else changes the heap's bookkeeping meanwhile, and starts updating its free lists.
1 / 5

?Why is the list of safe calls inside a handler so short?

Because the handler may have interrupted the very function it wants to call, midway through updating that function's state. So POSIX, the standard that Unix-like systems follow, publishes a short list of async-signal-safe functions, ones written so that calling them at any moment is harmless. It excludes malloc, printf and almost everything else. It does include write, the very call this chapter has been timing, so a handler can log a fixed message with one write() and no buffer.

The usual pattern is to do almost nothing in the handler. It sets a flag, declared volatile sig_atomic_t (a type the standard guarantees can be set in one uninterruptible step), or writes one byte into a pipe, a pair of descriptors where bytes written to one end can be read from the other. Then the main loop notices the flag or the byte and does the real work, which is a safe place to call printf. A pipe the program uses to talk to itself this way is called a self-pipe.

That's every way across the border, in both directions. What's left is finding out how often a real program crosses, and what to do about it.

08Counting your crossings

8.1Finding the crossings

Each question this chapter raised has a tool that answers it on a running program.

Shell
# How many crossings, and which calls? (sections 3 and 5) Start here.
strace -c -f ./app              # Linux
dtruss -c ./app                 # macOS (needs System Integrity Protection relaxed)
 
# Which call site made them? (section 5)
perf trace ./app                # Linux
perf record -e raw_syscalls:sys_enter -g ./app
 
# How much time goes to the kernel rather than your code? (section 3)
/usr/bin/time -v ./app 2>&1 | grep -E 'System time|User time'   # Linux; macOS: /usr/bin/time -l
 
# Is the interrupt load the problem? (section 6, Linux)
cat /proc/interrupts
mpstat -P ALL 1                 # %irq and %soft columns

strace -c is the first command to run, and most of the time it answers the question. A count of 4 million write calls where you expected a few thousand is the entire diagnosis.

8.2Reducing crossings, in order

  1. Buffer. The biggest win for the least effort, with no new API (section 5.1).
  2. Batch. writev, readv, sendmmsg and recvmmsg make one crossing carry many items.
  3. Avoid the copy. sendfile, splice, or mmap (which maps a file into your program's memory) for large file reads, so the data doesn't travel up into your program and back down.
  4. io_uring, on Linux 5.10 or newer, when you need hundreds of thousands of operations per second. It changes the shape of your code, so it belongs last.

8.3What you trade for what

You getYou payWhen the bill arrives
A protected kernel that checks everything~350 ns per crossingIn any loop that calls per item
vDSO and cached values for cheap callsYou can't tell which is which by reading CWhen you benchmark getpid and measure nothing
Buffered I/O, far fewer crossingsData sits in your program's memory until flushedOn a crash, as lost log lines
io_uring, potentially zero syscallsA harder API and a newer kernelAs complexity, up front
Spectre and Meltdown mitigations keeping you safeEvery crossing got more expensive from 2018 on affected processorsAs old benchmarks that won't reproduce

8.4Symptom, cause, fix

SymptomLikely causeFix
High system time you didn't expectA syscall per item: unbuffered logging, byte-at-a-time readsstrace -c -f to find it, then buffer
A syscall benchmark reads ~2 nsThe call is cached in libc and never crosses (getpid on macOS)Benchmark write to /dev/null instead
An old benchmark won't reproduceKPTI and Spectre mitigations since 2018Compare against a post-2018 baseline
High %irq or %soft in mpstatInterrupt load starving user threadsCheck /proc/interrupts; see NAPI and chapter 10
Process hangs inside a signal handlerA call that isn't async-signal-safe, like malloc or printfSet a flag or write to a self-pipe

09Summary

  1. User code can't touch hardware. Those instructions fault in user mode, so every disk and network request goes through the kernel, as buy milk did in section 3.
  2. There's one door in. A trap instruction jumps to a single entry point, and a number in a register picks the service.
  3. The kernel trusts nothing you pass. It checks every pointer and copies the data across, which is where much of the cost comes from.
  4. A real crossing costs about 350 ns, even for a write to /dev/null that does no work. A million of them is a third of a second.
  5. Some calls never cross. clock_gettime reads the vDSO or commpage in about 13 ns, and getpid on macOS returns a cached value in about 2 ns.
  6. You can't tell which is which from the C. Check with strace -c on Linux or dtruss -c on macOS.
  7. Mitigations made every crossing dearer in 2018. On affected processors, KPTI roughly doubled syscall cost on some workloads.
  8. Fast code crosses less. Buffering forty log lines into one write cut 17.4% of a core to about 0.4%. After buffering come batching, avoiding the copy, and io_uring, in that order.
  9. Interrupts use the same door from the other side. A short top half acknowledges the device and the bottom half does the work with interrupts enabled.
  10. Signal handlers may call almost nothing. Set a flag or write to a self-pipe, and do the work in the main loop.

10Build this

Measure your own border, and find a call that isn't crossing it. You already ran the first half of this in sections 1 and 4.

  • Time getpid, clock_gettime, write to /dev/null, and an empty loop.
  • Confirm with strace -c (Linux) or dtruss -c (macOS) which ones produce a real syscall. At least one will surprise you.
  • Then write one megabyte to a file, one byte at a time, and again in 4 KB blocks. Compare wall time and the strace -c counts.

Your 4 KB version should make about four thousand times fewer crossings (a million against 245) and run much faster. Having produced that ratio yourself makes buffering an instinct instead of advice.

11Interview questions

beginnerRoughly what does a syscall cost, and why isn't it free?›

A few hundred nanoseconds: about 350 for a write to /dev/null on a recent laptop, with the exact figure depending on the processor, the operating system and the security mitigations. The cost comes from the controlled crossing: switch privilege level, change stacks, dispatch through a table, check every pointer the caller passed, do the work, then restore everything. Checking matters because user code could change a pointer concurrently, so the kernel copies instead of trusting addresses.

intermediateWhy is calling clock_gettime so much cheaper than calling write?›

It usually isn't a syscall. Linux and macOS both map a small read-only page into every process (the vDSO on Linux, the commpage on macOS) containing code and a timestamp the kernel keeps current. Calling clock_gettime jumps into that page and reads the value without ever trapping.

That's about 13 ns against about 350 for write, roughly twenty-six times cheaper. The C looks the same either way, but one call crosses the border and the other doesn't.

intermediateYour service spends 20% of CPU in the kernel and you expected almost none. First step?›

strace -c -f. It gives you syscall counts and time by type in one command, and the answer is usually obvious: a few million write calls from unbuffered logging, or a clock read per loop iteration on a virtual machine where the clock can't be read from the shared page.

At roughly 350 ns each, half a million log lines a second is about 17% of a core spent purely crossing. Buffering to 4 KB blocks cuts that by a factor of forty.

deepWhat may you safely call inside a signal handler, and why is the list so short?›

Only async-signal-safe functions, a short POSIX list. A handler runs at an arbitrary instruction on the interrupted thread's stack, possibly while the interrupted code holds a lock or is midway through updating allocator state.

Call malloc from a handler that interrupted malloc and you corrupt the heap or deadlock. The safe pattern is to set a volatile sig_atomic_t flag or write one byte to a self-pipe, then do the real work from the main loop.

deepHow does io_uring reduce syscall cost to nearly zero?›

By replacing per-operation crossings with shared memory. Two ring buffers, submission and completion, are mapped into both the process and the kernel. The application writes submission entries directly into memory the kernel can already see and advances a tail pointer.

With IORING_SETUP_SQPOLL a kernel thread polls that ring, so the data path can run with no syscalls at all. You trade a much harder API and a dedicated polling thread for removing the border from the hot path entirely.

12Go deeper

check yourself
You benchmark syscall overhead with getpid and measure 2 ns. What went wrong?›

On macOS, and on Linux with older glibc, getpid isn't a syscall: libc caches the value at process start. Use something that must reach the kernel, like write to /dev/null, and confirm with strace -c or dtruss -c.

Why can't you call printf in a signal handler?›

It isn't async-signal-safe. The handler may have interrupted stdio or the allocator mid-update, so calling back in can deadlock or corrupt state.

Unbuffered logging at 500k lines/sec. Roughly what does it cost?›

About 17% of a core in crossings alone, at ~350 ns each. Buffering to 4 KB cuts it to under half a percent.

Why did syscalls get slower in 2018?›

Spectre and Meltdown mitigations: page table isolation and predictor flushing across the border. Old benchmarks stopped reproducing.

Operating Systems: Three Easy Pieces, chapter 6

"Mechanism: Limited Direct Execution" builds the user mode, kernel mode and trap instruction from this chapter one step at a time. Free online at ostep.org.

Jens Axboe: Efficient IO with io_uring

Design document from the author. Short, and the clearest explanation of the ring protocol.

Linux vDSO(7)

What's in the mapped page and why. Explains the 13 ns row in section 4.

signal-safety(7)

Actual list of what you may call in a handler. Shorter than you expect.

Brendan Gregg: perf trace

Lower-overhead than strace and usable on production traffic, which strace mostly is not.

The Linux Networking Stack

Interrupts, softirqs and NAPI at full depth, following one packet up to your read(). Chapter 10.

Speculative Execution, Spectre & the Mitigation Tax

Why crossings got more expensive in 2018, and what KPTI does. Chapter 44.

Filesystems & the Page Cache

What happens after read() and write() cross the border. Chapter 08.

eBPF: Running Your Code in the Kernel

Tracing syscalls on a live system with less overhead than strace. Chapter 48.