Here are two lines from a function that guards a small table. They read table[x] only when x is smaller than the table's length:
uint8_t table[16];
if (x < table_len) // the guard
y = table[x];If someone passes x = 4000, the guard fails and the read never happens. You can prove that from the code alone, because no path through these lines touches table[4000]. Now picture what sits 4000 bytes past the start of table: some other part of the program's memory, holding a byte you'd rather nobody read. Say its value is 83.
The processor doesn't take your proof as literally as you do. To evaluate the guard it first needs table_len, and if that number has to come from main memory it takes a long time to arrive. A core is one of the independent engines inside a processor that runs a stream of instructions, and a core that sat idle for that long would waste most of its life, so it guesses how the guard will turn out and starts running the code beyond it. When table_len arrives and the guess was wrong, the core throws the work away and the program never sees the result. For x = 4000, though, the thrown-away work included a read of the byte at table[4000], and that read left a trace in a place the core never cleans up.
In 2018 researchers showed that a patient attacker can read that trace and recover the byte, and the family of attacks that followed is called Spectre. This chapter asks one question: when a processor runs code the program said shouldn't run and then undoes it, can anything still leak, and what does stopping the leak cost every program on the machine? We'll follow that one byte from the core's guess, through the attack, to the defences, and finish by asking your own machine which of these attacks it's exposed to.
01Why a core guesses
1.1Waiting for one number
To check the guard, the core needs two numbers: x and table_len. Suppose table_len has to come from somewhere slow. How slow is that?
A core keeps small, fast copies of the memory it used recently in a cache, right next to it. The nearest and smallest cache is the L1 cache. Behind it sit larger and slower ones, such as the L2, and behind those is main memory, usually called DRAM. A load (a read of memory) that finds its data in L1 is an L1 hit, and one that has to go all the way to DRAM is a miss. Caches never copy a single byte. They copy the whole fixed-size chunk of memory around it, 64 bytes on most Intel and AMD chips and 128 on Apple's, and that chunk is called a cache line. Once one byte of a line has been loaded, the rest of the line is in the cache too, which matters later in this chapter. Chapter 02 builds this hierarchy properly.
To see how different the levels are, you can time a pointer chase: a program that walks a randomly shuffled chain of memory locations, where the address of each load is the value the previous load returned. The core can't overlap these loads, because it doesn't know the next address until the last load finishes, so the time per load is the latency of a single one. Making the chain cover more memory pushes it out of one cache after another. These numbers come from a recent Apple M-series core (Apple's own laptop chips). Other processors differ in the details and not in the shape.
| Memory the chain covers | ns per dependent load | Where it lives |
|---|---|---|
| 16 KB | 1.02 | L1 data cache |
| 1 MB | 5.32 | L2 |
| 8 MB | 8.26 | L2, near its edge |
| 256 MB | 93.60 | DRAM |
Compare the first and last rows. A load that hits L1 takes about a nanosecond, and one that goes to DRAM takes about 93, roughly 90 times longer. In 93 ns a core can get through hundreds of instructions. If table_len has fallen out of the cache, waiting for it before doing anything else would leave the core idle for all that time, once for every guard like ours. We'll come back to this 90-to-1 ratio in section 3, where it turns out to matter for a different reason. For now it explains why the core doesn't wait for table_len and guesses instead.
1.2Guessing which way the guard goes
The guard is a branch, a place where the program goes one of two ways depending on a value. The core has a unit called the branch predictor that guesses which way each branch will go before the answer is known, using what that branch did in the past. If our guard has passed the last thousand times, the predictor guesses that it will pass again, and the core starts fetching and running the instructions inside the if straight away.
How good can that guessing get? Call the two ways a branch can go taken (it jumps somewhere else) and not taken (it falls through to the next line), the names chip designers use. A predictor that only tracked which way a branch goes more often would be wrong half the time on a branch that strictly alternates, taken then not taken then taken. Real predictors keep history. To test that, a program fills a buffer with about a million (220) branch outcomes by repeating a pattern of a fixed length, called its period, over and over. Both sides of the branch do identical work, so the only thing that changes with the period is how often the predictor guesses wrong. The figure shows the time per branch as the period grows.
Strict alternation is a 50/50 branch with no bias at all, and it costs only 0.746 ns. The predictor has seen taken, not taken, taken, not taken, and worked out the rest. Push the period to 1024 and the cost is still 1.227 ns, well under the 3.239 ns of outcomes with no pattern to learn. So the predictor is a memory of recent behaviour, big enough to learn a repeating sequence a thousand branches long.
Guessing is a good bet most of the time, which is why every fast core does it. The price shows up on the occasions the guess is wrong, because by then the core has run instructions the program never asked for. What happens to them?
02Running ahead, then taking it back
2.1Holding the work until it's real
Taking the work back is possible because a modern core doesn't make its results official as it goes. It fetches instructions far ahead, runs each one as soon as its inputs are ready (which isn't always program order), and parks every result in a queue called the reorder buffer, or ROB. The ROB keeps the instructions in program order. A result only becomes official when its instruction reaches the front of the queue and is retired: registers (the core's few named storage slots) are updated, writes to memory are released, and any error the instruction caused is reported. On Apple's M-series cores the ROB is hundreds of entries deep, so hundreds of instructions can be in flight at once, none of them official yet. Running instructions on a guess, before it's known that they should run, is called speculative execution.

?What happens when the guess was wrong?
When the branch finally resolves and the guess turns out wrong, the core squashes everything it fetched after the branch. Those entries leave the ROB without retiring, the registers keep their old values, and no error is reported even if one of those instructions would have caused a fault, the error a core raises when an instruction does something forbidden, such as reading memory it isn't allowed to. As far as the program can tell, they never ran.
Here is our guard with x = 4000 and a slow table_len, step by step. Watch the register y and the cache in the bottom row.
x = 4000. table_len isn't in the cache, so it's about 93 ns away and the core can't tell yet whether the guard passes. The predictor has watched this guard pass every time, so it guesses pass.Look at the last two frames together. The register y was never set, so your program's view of the world is exactly what it would have been if the guard had never been guessed. The cache is a different story.
2.2What the rollback leaves behind
Chip designers split the processor's state into two kinds. Architectural state is everything your program can name and see: its registers and the contents of memory. The squash restores it perfectly. Microarchitectural state is the processor's own internal bookkeeping, which no instruction can read directly: which lines are in the cache, what the branch predictor has learned. The squash doesn't touch it, because nothing in the specification says it has to.
Think of a waiter who knows your usual order and pours your coffee before you've said a word. When you ask for tea instead, the coffee goes down the drain and you never drink it, so as far as your meal is concerned nothing happened. But someone watching the counter can see that a cup was poured, and which one. The squash is the coffee going down the drain. The warm cache line is the used cup still sitting on the counter.
?Why didn't anyone treat that as a bug before 2018?
Because for correctness it isn't one. Everything a program can compute depends on architectural state, and that is restored exactly after a bad guess. Speculation was always understood to be invisible, and for what a program computes, it is.
A warm line isn't a secret yet, though. The attacker chose x = 4000, so they already know which address was touched. To turn a warm line into the byte 83, we need to see what an attacker can add.
03Turning a warm line into a message
3.1Reading the cache with a clock
A program can't ask the cache what it holds. It can time a load, though, and section 1 showed that a hit takes about 1 ns and a miss to DRAM about 93 ns. So to find out whether a line is warm, load from it and look at the clock: fast means it was already in the cache, slow means it wasn't. A way for information to leak out through something that was never meant as an output, here the time a load takes, is called a side channel.
That gives the attacker a way to see warmth. What it doesn't give is the value 83. A warm line says only that somebody touched that address, and the attacker picked the address. For the byte itself to show up, its value has to decide which line gets touched.
3.2Letting the secret pick the line
The wrong-path code can do more than read the byte. It can use it. If the code takes the byte it read and uses it as an index into a second array, then which part of that second array gets loaded depends on the byte, and the attacker can time that array afterwards. This is the shape of the victim, in the pattern every write-up uses:
uint8_t table[16]; // the public table, 16 bytes
uint8_t probe[256 * 512]; // one slot per value a byte can have, 512 bytes apart
void victim(size_t x) {
if (x < table_len) // the guard
temp &= probe[table[x] * 512]; // read a byte, use it as an index
}Say the wrong path reads table[4000], which is 83, and then loads probe[83 * 512]. Slot 83 of probe ends up in the cache and the other 255 slots don't. Which slot is warm is the secret. (The temp &= only stops the compiler from deleting the load.) The attacker can call victim, and can read and time probe, but can't read past the end of table directly, because the guard stops it. That fits a real setting: victim could be a function inside the engine that runs a web page's JavaScript, whose memory the page isn't allowed to read.
?Why multiply by 512?
So that each possible byte value lands on its own cache line, far enough from its neighbours that loading one doesn't pull the next one in. If the slots were adjacent, several byte values would share a line and the timing couldn't tell them apart. A byte can have 256 values, so the probe array has 256 slots.
3.3Making the guard guess wrong on demand
One piece is still missing. The wrong-path read only happens if the predictor guesses "pass" for an x that fails the guard, and only matters if the core is slow enough getting table_len for the guess to run ahead. The attacker controls both. They call victim with a valid x, say 3, many times, so the predictor learns that the guard always passes. Section 1.2 showed that predictors learn patterns far more complicated than that. Then they push table_len out of the cache so it will be slow to arrive (section 4.2 says how), and call victim(4000). The predictor, which remembers only history, guesses "pass". Training a predictor on purpose so that it guesses wrong at a moment you choose is called mistraining it.
?Why is trainability a security problem?
Because the predictor is shared state. It has no idea whose code is running or what the branch means. Code that runs just before the victim, including code the attacker controls, can leave it trained to guess wrong in exactly the place it chooses.
We now have all three pieces: a predictor that can be trained, a secret-dependent load that picks a probe slot, and a clock that tells warm from cold. Let's run them together on our one byte.
04One byte, start to finish
4.1The attack on table[4000]
Putting the three pieces together gives the first variant, called Spectre v1 (variant 1): a bounds check that the core skips speculatively. The guard is correct, and architecturally an out-of-bounds x never reads anything. The attack works in the window before the guard resolves. Step through it, and keep an eye on where the secret byte goes: it never lands in anything the attacker can read directly.
table[4000], which victim would never hand over. All 256 slots of probe are cold. Three of them are drawn here: slots 0, 83 and 255.4.2What the receiver needs
The last step asks a lot of the attacker. Before the attack, all 256 slots of probe have to be cold, and afterwards the attacker needs a clock that can tell a 1 ns hit from a 93 ns miss. On x86, the instruction set of most Intel and AMD processors, both are easy. clflush is an instruction any program can run to evict one line from the cache, and rdtsc reads a counter that ticks every cycle. A proof of concept is a small program written to show that an attack works, and the textbook one for Spectre, in Appendix C of the original paper by Paul Kocher and colleagues, is written for x86 for exactly that reason.
Other processors may not offer either. Before we try the attack on a laptop, try predicting what it would see.
A program on an Apple M-series laptop has a clock that ticks every 42 ns and no flush instruction. An L1 hit takes about 1 ns and a DRAM miss about 93 ns. What does a naive Spectre receiver see?
05Trying the textbook attack on a laptop
5.1Zero of forty bytes
The paper's proof of concept can be ported to a Mac with two substitutions. _mm_clflush becomes a big eviction loop that streams through a 24 MB buffer, and rdtsc becomes mach_absolute_time, macOS's timer call. The logic is the paper's: mistrain the branch, call with the out-of-bounds index, then time all 256 slots of the channel. The secret is a 40-byte string beginning with The. Each byte is tried 999 times, and a slot's score is how many of those 999 tries it looked like a cache hit.
[diag] of 256 channel lines, 256 scored >=900/999 -- no line stands out
byte 0: got 0xff score 999/999 weak (want 'T')
byte 1: got 0xff score 999/999 weak (want 'h')
byte 2: got 0xff score 999/999 weak (want 'e')
...
recovered 0/40 bytes correctlyThe output calls the probe slots "channel lines", since each slot is one cache line of the channel. The first line says every one of the 256 scored at least 900 out of 999, so the receiver believes all of them were cache hits on nearly every try. For each byte it then reports the best-scoring slot, and with all 256 tied it falls back on the last one, slot 255, which is 0xff in hexadecimal. The word weak is the program's own verdict that no slot stood out. Against a secret starting The, it recovered 0 of 40 bytes.
?Did speculation not happen?
It did. The timer just can't see the result:
What the receiver has to read is a gap of 1 ns against 93 ns, using a ruler whose marks are 42 ns apart. A DRAM miss spans barely two marks and an L1 hit spans none. On top of that, a normal macOS program has no cache-flush instruction, so evicting the probe lines means streaming that 24 MB buffer. That's slow enough that the timing probably washes out further, and lines pushed out this way often stop in a nearer cache and still look like hits. The x86 proof of concept works because rdtsc resolves single cycles and clflush is available to any program. Neither is true here.
5.2The blinding is a defence
Web browsers run other people's JavaScript all day, so after 2018 they made the page's own clock, performance.now(), coarse enough that a page can't see gaps of a few nanoseconds. The port in section 5.1 ran into the same wall that browsers built on purpose. They also turned off SharedArrayBuffer, a way for two threads to share memory.
?Why disable SharedArrayBuffer?
Because a second thread that does nothing but increment a counter in shared memory is a home-made nanosecond timer. Coarsening performance.now() would be pointless if a page could build its own clock. Getting SharedArrayBuffer back now requires the page to be isolated from other sites, with two HTTP headers called COOP and COEP.
The experiment reproduced the cause, speculation past a bounds check, and couldn't reproduce the read. It doesn't show that a patient attacker fails. It shows only that the easy version does.
Spectre v1 was only the first way found to steer speculation. Once researchers saw that it leaves footprints, they went looking for others.
06One idea, many variants
6.1The family, 2018 to 2023
Spectre v1 mistrains a conditional branch, a branch that goes one of two ways like our guard. Several rows below involve the kernel, the central part of the operating system, which runs with full access to the hardware and holds memory no ordinary program may read. Every later variant has the same shape: the core does something early, on a guess or ahead of a permission check, and the result leaves a trace. They differ in which guess is steered and which internal structure leaks. A fix that narrows what an attack can reach, without removing the underlying behaviour, is called a mitigation. Dates matter here, because the mitigation each one forced is still running on your machine today. The CVE numbers are the public identifiers given to each security flaw.
| Name | CVE | Public | What speculates wrongly |
|---|---|---|---|
| Meltdown (v3) | CVE-2017-5754 | Jan 2018 | A load the kernel would refuse (a fault) still hands kernel data to later instructions before the fault retires |
| Spectre v1 | CVE-2017-5753 | Jan 2018 | A mistrained bounds check is skipped speculatively |
| Spectre v2 | CVE-2017-5715 | Jan 2018 | An indirect branch, a jump whose target is read from a register or memory, is steered to code the attacker chose |
| Spectre v4 (SSB) | CVE-2018-3639 | May 2018 | A load runs before an earlier store to the same address resolves, and reads the stale value |
| L1TF / Foreshadow | CVE-2018-3615 | Aug 2018 | A load from memory the page tables (section 7.2) mark as not present still reads whatever the L1 cache holds for that address |
| MDS (RIDL etc.) | CVE-2018-12130 | May 2019 | Small buffers inside the core, which hold data on its way to or from the cache, leak it to other code running on the same core |
| Retbleed | CVE-2022-29900 | Jul 2022 | A return instruction takes its target from the branch predictor, which defeats the main Spectre v2 defence (section 7.3) |
| Downfall (GDS) | CVE-2022-40982 | Aug 2023 | Intel's gather instructions, which load many values into one wide register at once, briefly expose stale data that other programs left in those registers |
| Inception (SRSO) | CVE-2023-20569 | Aug 2023 | On AMD processors, the predictor is trained to imagine call instructions that aren't there, which fills the return stack buffer (the small stack of return addresses) with a target the attacker chose |
?Why do the mitigations keep piling up?
Because they don't retire when the next bug lands. They stack. A kernel from 2024 carries the fixes for all nine rows, and section 8 adds up what they cost.
The gap from 2019 to 2022 wasn't quiet either. Researchers were realising that every predictor and every internal buffer is fair game, and vendors were shipping microcode (updates to the processor's internal firmware) faster than the papers came out.
6.2The ones that name Apple silicon
Most of that table is Intel and AMD. Apple silicon has its own attacks, and they target structures newer than the branch predictor:
| Attack | When | What it abuses | Shown on |
|---|---|---|---|
| GoFetch | March 2024, no CVE | The data memory-dependent prefetcher (DMP): hardware that watches the values a program loads and, when one looks like a pointer, loads the memory it points to ahead of time | M1, M2, M3 |
| SLAP | January 2025, no CVE | The load address predictor | M2 and later, stealing data from Safari |
| FLOP | January 2025, no CVE | The load value predictor | M3 and later, from Safari and Chrome |
GoFetch feeds the DMP cryptographic key material shaped like a pointer, and the prefetch leaks the key through the cache. Its authors never tested the M4, which didn't exist yet, so where it stands on that chip isn't known. SLAP and FLOP use prediction mechanisms newer than branch prediction, doing the same trick one level down: guess a load's address or result, and speculate on the guess.
6.3Asking your own kernel
With the names in hand, you can ask your own machine which of these it's exposed to. The Linux kernel keeps its own assessment of every speculation attack it knows about in a directory of small files, one per attack. Reading them is a one-line loop. On a Mac, Docker Desktop runs Linux inside a small virtual machine on the laptop's ARM chip (ARM is the design family behind phones and Apple's laptops), so what you see is Linux's opinion of an Apple core.
The script moves into that directory, then for each file prints its name left-aligned in 24 columns (printf "%-24s %s\n") followed by the file's contents ($(cat $f)).
docker run --rm alpine sh -c '
cd /sys/devices/system/cpu/vulnerabilities
for f in *; do printf "%-24s %s\n" $f "$(cat $f)"; done'gather_data_sampling Not affected
itlb_multihit Not affected
l1tf Not affected
mds Not affected
meltdown Not affected
mmio_stale_data Not affected
reg_file_data_sampling Not affected
retbleed Not affected
spec_rstack_overflow Not affected
spec_store_bypass Vulnerable
spectre_v1 Mitigation: __user pointer sanitization
spectre_v2 Not affected
srbds Not affected
tsx_async_abort Not affectedThere are fourteen files, one per attack, and each ends up with one of three answers: Not affected, Mitigation: followed by what the kernel does about it, or Vulnerable. Almost every name in the list is one more way of reading leftover state from speculative work. Meltdown, L1TF, MDS, Downfall (gather_data_sampling) and Retbleed were design problems in Intel or AMD processors, so an ARM core reports Not affected. Two lines say something else. spectre_v1 says the kernel added a code-level fix, and spec_store_bypass, the Spectre v4 of our table, says Vulnerable, which means no fix is switched on. Linux added this directory in version 4.15, alongside the first of the fixes.
What is the kernel doing on the Mitigation: line, and why is it fine with Vulnerable on the other? To answer, we have to look at the fixes.
07The defences
7.1Patches around the edges
The obvious fix is to stop speculating. That would make every core much slower, because speculation is most of why a modern core is fast (section 1). So the fixes are patches around the edges, each one breaking one kind of guess or closing one route out of the processor. The kernel carries most of them, and this table lists the ones that show up in its report.
| Mitigation | Fixes | How it works |
|---|---|---|
| KPTI / PTI (kernel page-table isolation) | Meltdown | The kernel stops mapping most of itself into your process's page tables (the address-translation tables, explained below), so a speculative load can't reach kernel memory. In Linux 4.15, released 28 January 2018. |
| Retpoline | Spectre v2 | Invented by Paul Turner at Google: "return" plus "trampoline". Replaces an indirect branch with a return sequence that steers speculation into a dead pause loop. |
| IBRS / eIBRS | Spectre v2, Retbleed | Indirect Branch Restricted Speculation, a hardware control that stops less privileged code from steering the kernel's predictions of indirect jumps (section 7.3); eIBRS is the enhanced, always-on form. |
| STIBP | Cross-thread v2 | Single Thread Indirect Branch Predictors: stops one SMT sibling (one of the two hardware threads that share a core) from training the other's predictor. |
| SSBD | Spectre v4 | Speculative Store Bypass Disable: turns off the reordering behind v4. |
__user pointer sanitization | Spectre v1 | The kernel masks array indices and memory addresses that come from programs, so even on a wrong path they can't point outside the range they were checked against. |
The last row is what our own kernel report showed for spectre_v1: the kernel doesn't try to stop the guard from guessing, it limits what an untrusted index can reach on the wrong path. The first row is the one every program pays for, so we'll take it apart.
7.2KPTI: taking the kernel out of your address space
Meltdown works on processors that check permissions late. A program loads from a kernel address, the load is refused with a fault, but the refusal is only acted on when that instruction retires, so for a moment the data was available to the instructions after it. Those can use it to pick a probe slot, exactly like our attack.
Before 2018, every process's page tables, the tables the core uses to translate the addresses a program uses into real memory (chapter 04), also mapped all of the kernel, marked off-limits. That meant a system call, a request from a program to the kernel (chapter 07), didn't have to switch tables. KPTI changes that. The tables your code runs with now map your own memory and a small entry stub and nothing else of the kernel, so a speculative load of a kernel address finds nothing to read. The price is that every system call has to switch to a second set of tables, with the whole kernel mapped, and switch back.

Switching tables causes a second cost. The core keeps a small cache of recent address translations, the TLB, and its entries are only valid for the tables that produced them. Without help, the core throws them all away on every switch. Processors that support PCID (process-context identifiers) can tag each TLB entry with the set of tables it belongs to, and then the entries survive the switch. Here's one system call with KPTI on:
7.3Retpoline, and why it wasn't the end
Spectre v2 steers an indirect branch: a jump whose target address is read from a register or from memory instead of being fixed in the code, which is what a call through a function pointer is. The predictor guesses the target the way it guesses our guard, and an attacker can train it to guess the address of code of their choice. Retpoline rewrites each indirect branch as a return sequence that sends any speculation into a harmless loop, and the real jump happens once the target is known.
?Why wasn't retpoline the end of Spectre v2?
Because retpoline assumes the predictor for return instructions can be trusted. A return normally takes its target from a small separate stack of return addresses. Retbleed, published in 2022 by researchers at ETH Zürich, showed that on some Intel and AMD processors a return could still take its target from the branch predictor, which an attacker can train, so the trampoline didn't help. The mitigations that followed were expensive: ETH measured 14% overhead with the AMD patches and 39% with the Intel ones.
7.4Back to the two lines
Now we can read the interesting lines of our report. Mitigation: __user pointer sanitization is the last row of the table: Linux's code-level fix for Spectre v1. The Vulnerable line is the more surprising one.
?Why does spec_store_bypass say Vulnerable?
Vulnerable means only that no fix for Spectre v4 is switched on. The file doesn't say why. On an ARM processor, Linux can switch the fix on only if the processor offers a control bit for it or the firmware underneath the kernel offers to do the work, and a virtual machine like Docker's may pass neither through, which leaves the kernel nothing to switch on. Cost is the other common reason. The fix slows every load that follows a store, so on x86 Linux by default applies it only to programs that ask for it. Either way, the line is something to look into before you decide it matters. On a bare-metal Intel Xeon the same directory is a wall of Mitigation: lines, and every one of them is spending time.
7.5What a browser does
A kernel guards the boundaries between processes. A browser has a harder job, because it runs attacker JavaScript on purpose, all day, and inside one process. Its defence has several parts, and the second one is the blinding from section 5.
| Defence | What it does | What it costs |
|---|---|---|
| Site isolation | Chrome puts each site in its own process, so a Spectre read inside a tab only reaches that site's memory | Memory: a process per site |
| Coarse clocks | performance.now() is coarsened so a page can't see gaps of a few nanoseconds | Fine timing in pages that legitimately want it |
Gated SharedArrayBuffer | Off by default since 2018, because a counter thread is a home-made timer | Getting it back needs the COOP and COEP headers, which promise the browser that the page shares no window and loads nothing from other sites without their consent, so it can be walled off on its own |
Every defence in this section adds work somewhere. How much?
08What the defences cost
8.1The tax on every system call
The cost of KPTI lands on something nearly every program does constantly: system calls. The tax is per crossing into the kernel, which makes it easy to reason about and very uneven between programs.
On a laptop you can't flip the mitigations to compare. mitigations=off is a boot parameter, a setting given to the kernel as it starts, and Docker Desktop's Linux boots inside a virtual machine whose start-up isn't under your control. What you can time there is the crossing itself. getpid does almost no work inside the kernel, so its cost is nearly all crossing, and a bare getpid on that Linux VM costs about 132 ns, averaged over 5 million calls of the actual SYS_getpid trap and not the C library's wrapper, since older C libraries cached the answer and never entered the kernel at all. The cost of the mitigations has to come from published measurements. For the cost of the mitigations themselves, the best source is still Brendan Gregg's 2018 analysis for Netflix.
| Number | Overhead | Source |
|---|---|---|
| `getpid` system call, Linux VM on an Apple laptop | 132 ns (baseline) | 5 million calls of the real trap |
| KPTI at about 50,000 syscalls/sec per CPU | ~2% | Gregg, Netflix, 2018 |
| KPTI across Netflix's production range | 0.1–6% | Gregg: depends on syscall rate |
| KPTI + PCID, 100 MB working set | 2.1% → 0.5% | Gregg: PCID cuts the TLB cost |
| Retbleed mitigations | 14% (AMD patches) to 39% (Intel patches) | ETH Zürich, 2022 |
?Why does the overhead range so widely?
Because the tax is paid once per crossing, so it tracks the syscall rate. Gregg expected "between 0.1% and 6% overhead with KPTI" across Netflix's fleet because their services make very different numbers of system calls, and he noted that at 50k syscalls/sec per CPU the overhead may be 2% and climbs as the rate increases. PCID softens the worst of it by letting the TLB survive the table swap, which is why his 2.1% case falls to 0.5% once it's on.
8.2Turning the tax into CPU time
Gregg's rule of thumb gives a cost per crossing. Two percent of every second is 20 ms, spread over 50,000 system calls, which is about 400 ns each. Chapter 07 timed a write() to /dev/null at about 348 ns on a recent laptop, and showed that a service making half a million of those calls a second already spends about 17% of a core on crossing the boundary. Now add the mitigation tax on a hot core at 300,000 system calls a second, the kind of rate a proxy or a database reaches:
| Gregg's rule of thumb | 2% of a second at 50,000 syscalls/s | 20 ms |
| Implied cost per crossing | 20 ms ÷ 50,000 | ≈ 400 ns |
| A hot core at 300,000 syscalls/s | 300,000 × 400 ns | 120 ms/s |
| Share of that core spent on Meltdown's fix, if the cost keeps scaling with rate | ≈ 12% | |
This is an extrapolation. Gregg's measurements reached about 6%, so a rate this high is beyond what he reported, and the real figure depends on the workload and on whether PCID is on. But a tenth of a core per hot core stops sounding small once it's multiplied by ten thousand cores and a power bill, and that's the calculation behind every "should we turn mitigations off" argument.
09What to do about it
There are two levers: turn the mitigations off, or make the taxed path shorter. Most teams should leave the first alone and reach for the second.
9.1When mitigations=off is defensible
mitigations=off is a Linux boot flag that disables every optional CPU mitigation at once: no page-table isolation, no retpoline, no clearing of internal buffers. It's a real lever with a narrow use. The deciding question is whether anyone untrusted can run code on the same hardware.
| Situation | mitigations=off? | Why |
|---|---|---|
| Single-tenant box, no untrusted code, ever | Defensible | No attacker shares the hardware; the leak has no reader |
| HPC (high-performance computing) or batch node you own end to end | Often yes | Trusted code only, and the tax hits tight syscall loops hard |
| Anything running other people's code | No | That's the whole threat model: a tenant (a customer sharing your hardware), an AWS Lambda function, a browser tab |
| Kubernetes multi-tenant nodes | No | A pod is untrusted code by definition |
| A laptop with a browser | No | JavaScript is untrusted code you invite in hourly |
?So why not flip it everywhere and reclaim the tax?
Because most hardware isn't single-tenant. On a box running one trusted workload, the mitigations defend against a threat that can't reach you, and turning them off is a rational way to reclaim that cost, which runs from a couple of percent up to the 14–39% of Retbleed's fix. A system that runs other people's code in one process, like Cloudflare Workers, has to treat this as part of its design, and a boot flag can't settle it.
9.2Make the boundary rare instead
The tax is paid per crossing, so the highest-leverage move is to cross less. Everything chapter 07 says about reducing system calls pays double once the mitigations are in the price. Ranked by what it returns:
- Buffer. One
writeper 4 KB instead of one per line. Forty crossings become one, and each crossing now carries the KPTI tax too, so the win is bigger than it was in 2017. - Batch.
writev,sendmmsgandrecvmmsgtake many buffers or messages in a single system call. - Share memory with the kernel. io_uring, a Linux interface, moves submission and completion into queues in memory that both sides read, so a busy server can run its data path with almost no system calls. A path you don't cross pays no tax.
- Keep PCID on. It's on by default on anything modern. On a custom kernel, don't disable it to "simplify", because it's most of what makes KPTI affordable.
Batching and PCID are the two people forget. A system call you didn't make is the only one that's reliably free, mitigations or not.
10Checking a real machine
10.1Commands
Each question this chapter raised has a command that answers it on a Linux machine.
# Which attacks is this kernel exposed to, and what does it do about each? (section 6.3)
grep -H . /sys/devices/system/cpu/vulnerabilities/*
lscpu | grep -i vulnerability # the same information, formatted
# Was a boot flag changed? (section 9.1)
cat /proc/cmdline # look for mitigations=
# Is page-table isolation on, and does the CPU have PCID? (section 7.2, x86)
dmesg | grep -i isolation # may need root
grep -o -w -E 'pcid|invpcid' /proc/cpuinfo | sort -u
# How many system calls does this process make per second? (section 8)
sudo perf stat -e raw_syscalls:sys_enter -p $PID -- sleep 10
sudo strace -c -f -p $PID # per-call counts; slows the process a lotThe count from perf stat, divided by ten seconds and then by the number of CPUs the process keeps busy, is the number to hold against Gregg's 50,000 per second per CPU: far below it and KPTI is a rounding error for this process, far above it and the tax is worth looking at.
10.2Rules that hold up
- Read the kernel's own report before guessing what a machine is exposed to. A
Vulnerableline is a question to answer (no fix offered, or a fix left off on purpose), and aNot affectedline is a fact about the processor. - Leave mitigations on wherever other people's code can run: containers, shared hosts, browsers.
- Make crossings rare. Buffer, batch, and share memory with the kernel. This repays the tax whether or not you ever touch a flag.
- Keep PCID on. In Gregg's test it was the difference between 2.1% and 0.5%.
- Never read a failed proof of concept as proof of safety. Compare the timer's resolution with the hit/miss gap first.
10.3What you trade for what
| You get | You pay | When the bill arrives |
|---|---|---|
| Kernel memory unreachable to speculative loads (KPTI) | A table switch on every system call, 0.1–6% at Netflix's rates | As CPU time on syscall-heavy services |
| Retbleed protection | 14% to 39% on the affected processors | As a slowdown after a kernel update |
| Coarse browser clocks and isolated sites | Memory per site, no fine timing | As a browser using more RAM |
| Mitigations switched off | No protection against anyone sharing the hardware | As a breach on shared hosts |
10.4Symptom, cause, fix
| Symptom | Likely cause | Fix |
|---|---|---|
| A pre-2018 syscall benchmark won't reproduce | KPTI on every crossing | Check PCID is on; buffer and batch system calls |
| Kernel-heavy workload 14–39% slower after a kernel update on Intel or AMD | Retbleed mitigation | Accept it on shared hardware; consider mitigations=off only on single-tenant boxes |
vulnerabilities/* shows Vulnerable | No fix is switched on: the hardware or firmware offers none, or the kernel's default leaves it off | Find out which, then decide whether untrusted code ever shares the host |
SharedArrayBuffer is undefined in your page | The page isn't cross-origin isolated | Serve the COOP and COEP headers |
| A Spectre proof of concept scores every probe line as a hit | Timer too coarse, or no flush instruction | Compare the clock's resolution with the hit/miss gap |
11Summary
- A core guesses past a slow check. A DRAM load takes about 93 ns against about 1 ns for L1, and the branch predictor learns repeating patterns past period 1024, so guesses are usually right and the core rarely waits.
- Wrong guesses are held in the reorder buffer and squashed. Registers keep their old values and no fault is reported, so the program never officially ran the work.
- The squash doesn't empty the cache. Architectural state is restored exactly and microarchitectural state isn't, which is the whole of Spectre's insight.
- A clock turns warm and cold into information. One nanosecond against about 93 is a side channel for anything whose address depends on a secret.
- Spectre v1 turns a skipped bounds check into a read. The secret byte picks which of 256 probe slots gets loaded, and timing the slots recovers it, after the attacker trains the predictor and pushes the length out of the cache.
- Coarse clocks blind the reader and don't fix the leak. A naive port against a 42 ns timer recovered 0 of 40 bytes, and browsers ship that blindness on purpose.
- The variants keep coming and the mitigations stack. Apple silicon has its own: GoFetch, SLAP and FLOP abuse the prefetcher and load predictors.
- Each defence breaks one guess. KPTI unmaps the kernel, retpoline redirects indirect branches, and Retbleed showed retpoline's assumption was wrong, at 14% to 39%.
- KPTI's cost tracks the syscall rate. Netflix saw 0.1–6%, and PCID cut one case from 2.1% to 0.5%.
mitigations=offis only safe where no untrusted code shares the hardware. Pods, Lambdas and browser tabs all count as untrusted.- Crossing the boundary less beats removing the tax. Buffer, batch, use io_uring, and keep PCID on.
12Build this
Repeat the failed attack from section 5, and learn from why it fails.
- Take the Kocher Appendix C proof of concept and compile it on an Apple Mac with
clang++ -O2. Swap_mm_clflushfor a big eviction loop andrdtscformach_absolute_time. - Run it. Expect 0 bytes and every channel line scoring alike, which is the timer wall from section 5.1 and says nothing yet about whether speculation happened.
- Prove the wall is real. Read
mach_absolute_timetwice back to back in a tight loop and look at the smallest non-zero difference. You'll find one tick, about 42 ns, so a DRAM miss is only about two ticks and a miss that stops in a nearer cache is less than one. - Now watch the cause work anyway: a randomised dependent pointer chase across 16 KB versus 256 MB gives you the 1 ns / 93 ns gap the attack depends on. The ingredient is there, and only the reader is blind.
What's left to try is the counter-thread timer, amplified over many trials the way the Apple-silicon papers do it. A counter thread is a second thread that does nothing but increment a shared counter, and its value is read as a clock. A naive one runs slowly and unevenly, because every increment has to pass the counter's cache line between two cores, and macOS offers no way to pin the thread to one core, so it's harder than it sounds. Whether a patient attacker recovers the byte on a given Mac is an open question.
13Where you meet this in the wild
Boru Chen and colleagues showed Apple's data memory-dependent prefetcher leaking cryptographic keys on M1, M2 and M3 by disguising key bytes as pointers. M3 lets software set the DIT bit to disable the prefetcher; M1 and M2 don't expose the switch, so the mitigation lands in the crypto library instead.
Georgia Tech's team abused the load-address predictor (M2+) and the load-value predictor (M3+) to read data out of Safari and Chrome. The predictors guess a load's address or result and speculate on the guess, exactly the mistrainable structure Spectre needs.
Brendan Gregg measured Meltdown's fix across production and found overhead tracking syscall rate: 0.1% to 6%, about 2% at 50k syscalls/sec per CPU. The write-up is still the clearest public account of the tax landing on real traffic.
After Spectre, Chrome moved every site into its own renderer process and coarsened its timers. It's expensive in memory, and it shipped anyway, because the browser runs hostile code by design and can't trust the CPU to keep secrets across a boundary inside one process.
14Interview questions
beginnerSpeculation is supposed to be invisible. How does Spectre see it?›
Architectural state, the registers and memory a program can name, is restored perfectly on a bad guess. Microarchitectural state isn't rolled back. A cache line pulled in during the speculative detour stays warm after the squash.
Spectre reads that warmth. The secret never appears in any register the attacker can name. It appears as which cache line is fast, and a timer recovers it. The rollback is complete for everything the CPU promised to undo, and cache occupancy was never on that list.
beginnerWhat's the difference between Meltdown and Spectre v1?›
Meltdown (CVE-2017-5754) reads kernel memory from user space: a faulting load grabs privileged data before the fault retires. Its fix is KPTI, which unmaps the kernel from user page tables.
Spectre v1 (CVE-2017-5753) stays within one privilege level and mistrains a bounds check so the CPU speculatively reads past an array. No page-table trick stops it, and the fix is to mask untrusted indices so that even a wrong path can't use them to reach outside the array. Meltdown mostly hit Intel, while Spectre v1 hits nearly every speculating core.
intermediateWhy did syscalls get slower in 2018, and by how much?›
KPTI. Meltdown's fix stops mapping the kernel into your process, so every syscall swaps page tables entering and leaving the kernel, and that costs TLB work on top of the trap.
Brendan Gregg expected 0.1% to 6% at Netflix, roughly 2% at 50k syscalls/sec per CPU, scaling with syscall rate. PCID cuts it substantially (his 2.1% case fell to 0.5% with PCID on) because the TLB no longer flushes on every crossing.
intermediateYou ported the textbook Spectre PoC to a Mac and it recovered nothing. Bug or defence?›
Defence, mostly. The out-of-bounds read still happens speculatively, and the problem is reading the cache footprint back. macOS gives user code a timer that ticks every 42 ns and no flush instruction. Even a full DRAM miss is only about two ticks of that timer, and without a flush the probe lines probably never get evicted that far, so most "misses" land in a nearer cache and read the same as a hit.
That's the coarse-timer mitigation browsers ship, met from the attacker's side. A serious attack builds a finer timer (a counter thread) and averages many trials. The naive port doesn't, so it's blind.
deepWhen is mitigations=off a reasonable choice?›
When no untrusted code shares the hardware. A single-tenant box running one trusted workload, or an HPC node you own end to end, gets no protection from mitigations against an attacker who can't reach it, so reclaiming their cost (a couple of percent for KPTI, up to 14–39% for Retbleed's fix) is rational.
It stops being reasonable the instant someone else's code runs on that silicon: a Kubernetes pod, a Lambda, a browser tab. There the flag turns a hardware bug into a data breach. The threat model is "shared hardware", and the flag is only safe where that's false.
deepRetpoline mitigated Spectre v2. Why did Retbleed need more?›
Retpoline replaces indirect branches with a return sequence that steers speculation into a safe loop. It assumes the return predictor is trustworthy.
Retbleed (2022) broke that assumption: on some Intel and AMD parts, return instructions could take their speculative target from the branch predictor and not the return stack, so a poisoned predictor defeats the trampoline. The new mitigations are stronger and costlier than retpoline, and ETH Zürich measured 14% overhead with the AMD patches and 39% with the Intel ones. That cost is why they were contentious.
15Go deeper
A secret never lands in a register the attacker can read. How does the byte escape?›
As a cache line. The speculative load uses the secret to pick which line of a probe array to touch; that line stays warm after the squash, and its access time reveals the byte. The channel is timing and carries no data directly.
Your Spectre PoC scores every channel line as a hit. What's the first thing to check?›
Timer resolution. If a single hit and a single miss both fit inside one clock tick, the receiver is blind. On Apple laptops that tick is about 42 ns and a full DRAM miss is about 93 ns, right at the edge. Without a flush instruction the lines rarely get that far out, and a naive single-shot read loses.
Two identical single-tenant boxes; one has mitigations=off. Is that a bug?›
Not necessarily. With no untrusted code on the hardware, the mitigations defend against an attacker who can't reach it, so turning them off to reclaim syscall throughput is defensible. The same flag on a multi-tenant node is a breach waiting to happen.
Why did browsers disable SharedArrayBuffer after Spectre?›
Because a second thread incrementing a counter is a home-made high-resolution timer, and Spectre needs one. Coarsening performance.now() is pointless if you can build your own clock. COOP/COEP isolation is now required to get it back.
The original paper, and Appendix C is the proof of concept everyone ports. Read it next to the Meltdown paper; the split between architectural and microarchitectural state is the whole idea.
Jann Horn's 3 January 2018 write-up, the disclosure that started it. Longer and more concrete than the academic papers, and it shows the reasoning in the order he found it.
The 2018 Netflix analysis. Where the 0.1–6% figures come from, with the syscall-rate dependence and the PCID recovery laid out on real workloads.
The two Apple-silicon writeups: GoFetch's DMP attack and the SLAP/FLOP predictor attacks. The material for the machine you're most likely holding.
ETH Zürich's page for the 2022 return-instruction attack, including the 14% to 39% overhead of the AMD and Intel mitigations.
16Related chapters
Where branch prediction and out-of-order execution are the feature. This chapter is the bill for that speed. Chapter 01.
The path the mitigation tax lands on, and every technique for crossing it less. Chapter 07.
The cache timing gap the side channel is built from, measured properly. Chapter 02.
The page tables and TLB that KPTI swaps on every system call. Chapter 04.