Two threads, which are two independent streams of instructions inside one program, share a couple of variables. Thread 1 works out an answer and stores it in data, then sets a flag called ready to say the answer is there. Thread 2 waits until the flag is up and then reads data.
// thread 1 // thread 2
data = 42; while (!ready) {}
ready = true; print(data);Read it top to bottom and it can't go wrong. Thread 1 stores the answer before it raises the flag, and thread 2 only looks at the answer after it has seen the flag.
It can go wrong all the same. Run it enough times on the right machine and thread 2 will sometimes see ready as true and then read the old data, the 0 that was there before. Thread 1 did write in the right order. What matters is the order in which other cores see the writes, where a core is one of the independent units inside a processor, each running one thread at a time. Neither the compiler (the program that turns your source code into machine instructions) nor the processor promises to keep that order the same as the one you wrote. Section 4 runs this exact example and counts how often it breaks.
The rulebook that says which value a read may return when other threads are writing is called the memory model. This chapter asks one question about it: when one thread writes A and then B, when can another thread see B without A, and what do we write to forbid it? We'll start with a simpler way shared variables go wrong, find where the reordering comes from, fix the example above, look at the instructions the fix turns into, and finish with what it costs.
01Two threads and one counter
Reordering is the hard problem, but a shared variable can go wrong in a simpler way, and you can watch it happen in a few seconds.
1.1A million increments each
We'll start two threads that each add 1 to the same counter a million times, so the right final answer is 2,000,000. The program keeps two counters side by side: a plain long, and a std::atomic<long>, a type we'll explain right after the run. The volatile on the plain counter only stops the compiler from folding the whole loop into one addition of 2,000,000, which it would otherwise be free to do. (Section 6 shows that volatile does nothing else useful for threads.) std::thread starts a thread running the function it's given, and join waits for that thread to finish.
Save the program as race.cpp, build it with clang++ -std=c++20 -O2 race.cpp -o race (with g++, add -pthread), and run ./race. The -O2 flag turns on the optimiser, as in a real build.
#include <atomic>
#include <cstdio>
#include <thread>
volatile long plain = 0; // volatile only stops the compiler folding the loop
std::atomic<long> safe{0};
int main() {
const long n = 1000000;
auto work = [&] {
for (long i = 0; i < n; i++) {
plain = plain + 1; // read, add, write: three steps
safe.fetch_add(1); // one indivisible step
}
};
std::thread a(work), b(work);
a.join(); b.join();
std::printf("expected %ld\nplain %ld\natomic %ld\n", 2 * n, (long)plain, safe.load());
}expected 2000000
plain 1158128
atomic 2000000The totals vary from run to run. In three runs the plain counter ended at 1,158,128, 1,205,109 and 1,396,534, between 58 and 70 percent of the right answer, so somewhere between 600,000 and 840,000 increments vanished. The atomic counter came out at exactly 2,000,000 every time.
1.2Why the plain counter loses updates
plain = plain + 1 looks like one step and takes three. The core reads the counter from memory into a register, one of the few tiny storage slots inside the core itself, adds 1 there, and writes the result back to memory. A read from memory into a register is called a load. A write from a register back to memory is called a store, and we'll use both words for the rest of the chapter.
Two threads can interleave the load, the add and the store. Suppose the counter holds 5:
| Step | Thread A | Thread B | Counter in memory |
|---|---|---|---|
| 1 | reads 5 | 5 | |
| 2 | reads 5 | 5 | |
| 3 | adds 1, has 6 | 5 | |
| 4 | adds 1, has 6 | 5 | |
| 5 | writes 6 | 6 | |
| 6 | writes 6 | 6 |
Two increments ran and the counter went up by one. B loaded the counter before A's store reached memory, so B started from a stale 5, and its store of 6 overwrote A's work. Our two loops collide like this hundreds of thousands of times.
An atomic operation fixes this by making the load, the add and the store one indivisible step. (The word comes from the Greek for "cannot be cut".) No other thread can see the counter halfway through or slip a store in between. fetch_add(1) is the atomic version of plain + 1, and it returns the value the counter had before the add. An operation like this, which loads a value, changes it and stores it back as one step, is called a read-modify-write.
Each write in our opening example is complete, and nothing gets lost, so atomicity can't be what goes wrong there. The trouble is the order in which another thread sees the two writes. To find where that comes from, we need to look at what happens to our code between the source file and the silicon.
02Why your code doesn't run in the order you wrote it
We picture a program running line by line, each line finishing before the next one starts. For a single thread, that picture is never wrong in a way the thread could detect. It still doesn't describe what happens, because two separate layers rearrange your code on its way to running.
2.1The compiler moves code
The compiler rearranges your code whenever the result is the same for the thread that runs it. It keeps a variable in a register instead of writing it back to memory on every loop iteration, and it lifts a load out of a loop if nothing in the loop changes the variable.
Apply that to our example. Thread 1's two stores don't depend on each other, so the compiler may swap them. Thread 2 spins on ready, and nothing inside its loop writes ready, so the compiler may read it once before the loop and then spin forever on the stale copy. Each change is legal because the compiler's only obligation is that a single thread can't tell the difference. Nothing obliges it to consider what another thread might see.
2.2The processor parks stores in a buffer
The processor has its own version of this. Memory has a hierarchy of caches, small fast memories near each core. They hold copies of recent memory in cache lines, blocks of nearby bytes (typically 64 or 128) that move between memory and cache as a unit. When thread 1 writes data and the line holding data isn't in its core's cache, the core must fetch the line first. Chapter 02 put a fetch from main memory at about 83 ns, long enough to run hundreds of instructions, and no core is going to stall for that on every write.
So the core parks the store in a small private queue called the store buffer and carries on with the next instruction. The store waits there until the line arrives. Watch one store go through it:
data = 42. Memory holds data as 0, and the cache line that contains data isn't in this core's cache.?Why doesn't your own thread ever notice?
Because the core looks in its store buffer before it goes to memory, as frame 3 shows. From the inside, everything happens in program order. Only other threads, which read through the cache and never see the buffer, observe stores late, and possibly in a different order from the one you wrote.

One parked store hurts nobody. The trouble starts when thread 1 makes two stores in a row, the first has to wait in the buffer, and the second doesn't. That's exactly what our example does next.
03Publishing data with a flag
Now we put the store buffer to work on the opening example, with a second thread watching.
3.1Data, then a flag
Let's give the two stores in thread 1 names. A is data = 42 and B is ready = true. Thread 2 spins until ready is true and then reads data. This pattern, filling something in and then raising a flag so others know it's complete, is called publishing the data.
Whether it goes wrong depends partly on the processor, so we need names for the two big families. ARM is the family of processors in phones, in Apple's M-series Macs and in cloud servers such as AWS Graviton. x86-64, from Intel and AMD, is in most PCs and in most other servers. They reorder memory accesses by different amounts, and section 6 compares them. For now, try to reason the answer out from the store buffer alone:
The reader sees ready == true. Is it guaranteed to read the new value of data?
Here is how it goes wrong on ARM. Suppose ready's cache line is already in the writer's cache, and data's isn't:
data as 0 and ready as false. Thread 1 is about to run A, then B. The line that holds ready is already in the writer's cache, and the line that holds data isn't.The writer's store buffer is one of two places this can happen.
?Can the reader get things out of order too?
Yes. Thread 2's core is also allowed to run its load of data early, before its load of ready has finished, for instance by guessing that the spin loop is about to exit and reading ahead. A stale data can therefore come from either end, and the fix has to constrain both. This bug often turns up only after code moves to ARM servers or Apple Silicon laptops, months after it shipped, and section 6 explains why.
3.2What the language says about unsynchronised access
Is the worst case that a reader sees an old value? It's worse than that. Two threads touching one ordinary variable, with at least one of them writing and nothing ordering the accesses, is called a data race. In C++ a data race is undefined behaviour, which means the language puts no limit at all on what the program may do.
The compiler is entitled to assume that a valid program contains no data races, and it transforms the surrounding code on that basis. That's why racy programs can work for years and then break after a compiler upgrade with no change to the source. The plain counter in section 1 was a data race too, volatile or not. Lost increments were only the form of it that we could see.
The next section builds that rulebook from one relation.
04Happens-before, and five orderings
To forbid the bad outcome from section 3, we have to say that A must be visible to thread 2 by the time thread 2 reads data. The memory model has one relation for saying that.
4.1Happens-before
If action X happens-before action Y, then everything X wrote is visible to Y, meaning Y's loads return it. In our example we want A, the store to data, to happen-before thread 2's load of data. Two accesses to the same ordinary variable, at least one of them a store, must be ordered this way one way round or the other. If neither happens-before the other, that's the data race from section 3.2, and the program has no defined meaning.
Happens-before is assembled from two sources. The first is program order inside a thread: A happens-before B because thread 1 wrote them in that order. The second is a link between threads, made when an operation on an atomic variable in one thread is matched with an operation on the same variable in another. That link is called a synchronises-with edge. Atomic operations with the right ordering create those edges, and so do the tools built on top of them, such as mutexes and starting or joining a thread. Plain variables never create one.
Chain the two and you get the guarantee we want:
A: data = 42 ──program order──▶ B: store ready
│ synchronises-with
▼
C: load ready (sees B's value) ──program order──▶ D: read dataA happens-before B, B synchronises-with C, and C happens-before D, so A happens-before D, and D must see 42. What remains is how to ask for the middle edge.
4.2The orderings, and what each edge costs
Every atomic operation takes an ordering argument that says which edges it creates and which neighbouring accesses it keeps in place. C++ offers five that matter, and if you pass none you get the strongest, seq_cst (short for sequentially consistent). (A sixth, consume, is treated as acquire by every major compiler, so you can ignore it.)
| Ordering | Guarantees | Creates an edge? |
|---|---|---|
| relaxed | Atomicity only. No ordering with anything. | No |
| acquire (load) | Nothing after it moves before it | Yes, with a matching release |
| release (store) | Nothing before it moves after it | Yes, with a matching acquire |
| acq_rel | Both at once, on a read-modify-write such as fetch_add | Yes |
| seq_cst | All of the above, plus one order of every seq_cst operation that all threads agree on | Yes |
The atomic counter from section 1 needs only relaxed, because nothing else depends on its value. For our example, we store ready with release. That's a promise that everything this thread wrote before the store becomes visible no later than the store itself does, so A can't be left behind in the buffer. The reader loads ready with acquire, which does two things. If it sees the value the release stored, it is guaranteed to see everything written before that store. And none of the reader's later loads can be performed early, which closes the other half of the problem from section 3.
?What does an acquire do on its own?
Nothing. Acquire and release only work in pairs, on the same variable. An acquire load with no matching release store compiles perfectly and guarantees nothing.
4.3Fixing the example
Here is our opening example again, with ready made an std::atomic<bool> and ordered:
// thread 1
data = 42; // A: an ordinary variable
ready.store(true, std::memory_order_release); // B
// thread 2
while (!ready.load(std::memory_order_acquire)) {} // C
print(data); // D: guaranteed to print 42Notice that data itself can stay an ordinary variable. The edge between B and C orders the accesses to it, so there's no race. Now the same scene as before, with B as a release store:
How the hardware keeps a release store behind earlier stores differs between processors, and section 5 shows what it looks like on ARM. Whatever it does, the promise at the language level is the one in frame 3.
Before we run it for real, let's predict what we'll see.
Thread 1 stores data = 42 and then ready = 1, over two million rounds. Thread 2 counts the rounds in which it saw ready == 1 and data == 0. With relaxed orderings on ready, that count is above zero on ARM. What is it with a release store and an acquire load on ready?
The program below makes two million independent copies of the example. Each Round holds a data and a ready, and the 124 bytes of padding put ready 128 bytes after data, so the two never share a cache line, just as in the Scene. alignas(256) starts every round on its own 256-byte boundary, so neighbouring rounds don't share lines either and don't disturb each other. (Two million rounds of 256 bytes is about 512 MB of memory.) The writer stores data and then ready. The reader loads ready and then data, and counts a round as bad if it saw ready == 1 and data == 0. We make data an atomic accessed with relaxed so that the failing run is merely wrong and not undefined behaviour. The run is done twice, and only the ordering on ready changes between the two.
#include <atomic>
#include <cstdio>
#include <thread>
constexpr int N = 2000000;
struct alignas(256) Round { // each round gets its own data and ready,
std::atomic<int> data{0}; // 128 bytes apart, so they never share
char pad[124]; // a cache line (lines are 64 or 128 bytes)
std::atomic<int> ready{0};
};
static Round rounds[N];
// Writer: data = 42, then ready = 1. Reader: if ready is 1, look at data.
long run(std::memory_order store_order, std::memory_order load_order) {
for (auto& r : rounds) { r.data = 0; r.ready = 0; }
std::atomic<int> go{0};
long stale = 0;
std::thread writer([&] {
while (!go.load()) {}
for (auto& r : rounds) {
r.data.store(42, std::memory_order_relaxed); // A
r.ready.store(1, store_order); // B
}
});
std::thread reader([&] {
while (!go.load()) {}
for (auto& r : rounds) {
int ready = r.ready.load(load_order);
int data = r.data.load(std::memory_order_relaxed);
if (ready == 1 && data == 0) stale++; // saw the flag, missed the data
}
});
go = 1;
writer.join(); reader.join();
return stale;
}
int main() {
std::printf("relaxed store, relaxed load: %ld stale reads in %d rounds\n",
run(std::memory_order_relaxed, std::memory_order_relaxed), N);
std::printf("release store, acquire load: %ld stale reads in %d rounds\n",
run(std::memory_order_release, std::memory_order_acquire), N);
}relaxed store, relaxed load: 17894 stale reads in 2000000 rounds
release store, acquire load: 0 stale reads in 2000000 roundsLook at the first line. Each stale read is one round in which the reader saw the flag and missed the data, either because the writer's two stores became visible out of order, as in the second Scene, or because the reader's load of data ran early. The count changes from run to run, and on an ARM laptop it is around one round in a hundred. It only has to be above zero to prove the point. The second line is 0 every time, because the release and acquire pair rules both of those stories out.
One in a hundred sounds frequent, but this program does nothing except race the two threads against each other. In a real program the writer publishes once in a while and the reader checks at some unrelated moment, so the window is hit far more rarely, often rarely enough that a test suite passes for months and the bug turns up in production.
In real code, data is rarely a single integer. It's an object that took many writes to build, and ready is a pointer to it:
Config* cfg = new Config{...}; // 1. construct
ready.store(cfg, std::memory_order_release); // 2. publish
// another thread
Config* c = ready.load(std::memory_order_acquire);
if (c) use(c->field); // 3. safe to readThe release covers every field written in step 1, however many there are, in the same way it covered data in our example. A reader that loads a non-null pointer with acquire sees a fully built Config. Publishing a pointer this way is the most common real use of acquire and release.
4.4Where you already use it
The same pair sits inside code you already call. std::shared_ptr and Rust's Arc count how many owners an object has. Each owner that finishes decrements the count with a release, and the thread that takes it to zero does an acquire before freeing the object:
The Linux kernel spells the same pair smp_store_release and smp_load_acquire.
So far the orderings are promises written in source code. Something has to turn them into instructions that a processor can carry out, and that's the next question.
05What the orderings compile to
5.1One model for many processors
The compiler does the translation, and what it emits depends on the processor it's compiling for. That was the plan from the start. The memory model arrived in C++11, the 2011 revision of the language, and it was written to describe real hardware: it had to be implemented efficiently on x86, ARM, POWER and Itanium, which all reorder in different ways. Each of those processors has its own instructions for limiting reordering. Instead of exposing them, the C++ standards committee defined the weakest guarantees that are still useful, the orderings from section 4, and left each compiler to emit whatever its target needs to keep them.
?So does release mean "insert a fence"?
A fence, also called a barrier, is a separate instruction that stops the processor from reordering memory accesses across it. Release doesn't mean "insert a fence". It means "make this synchronises-with edge work", and the compiler may do that with a fence, with a special store instruction, or with nothing at all, depending on the target.
5.2On ARMv8, no barrier at all
A common picture says that acquire and release insert a barrier instruction. On ARMv8, the current generation of 64-bit ARM, the compiler's output shows something else. You can see that output with clang's -S flag, which prints the machine instructions it produced in readable form, called assembly. Here are eight tiny functions, each doing one atomic operation on a 64-bit value, compiled with -O2 -S for ARMv8.3 or later:
ld_rlx: ldr x0, [x8] ; relaxed load — an ordinary load
ld_acq: ldapr x0, [x8] ; acquire load — Load-Acquire RCpc
ld_sc: ldar x0, [x8] ; seq_cst load — Load-Acquire RCsc
st_rlx: str x0, [x8] ; relaxed store — an ordinary store
st_rel: stlr x0, [x8] ; release store — Store-Release
st_sc: stlr x0, [x8] ; seq_cst store — THE SAME INSTRUCTION
add_rlx: ldadd x9, x0, [x8] ; relaxed fetch_add
add_sc: ldaddal x9, x0, [x8] ; seq_cst fetch_addEach line is one function's name, then one instruction: its opcode, the name of the operation such as ldr (load register), followed by the register and the address it works on. Here is what to notice in those eight lines.
| What you see | What it means |
|---|---|
No dmb anywhere | dmb is ARM's separate barrier instruction. None appears: the ordering is encoded in the load and store opcodes themselves |
| seq_cst store = release store | Both are stlr, so the compiler has nothing extra to emit for the stronger ordering |
ldapr for acquire, ldar for seq_cst | Both are acquire loads. ldapr (ARMv8.3 and later) is the weaker and cheaper one, and it lets a later load overtake an earlier stlr. seq_cst forbids that, so it needs ldar. Older ARM targets use ldar for acquire too |
ldadd vs ldaddal | One atomic read-modify-write instruction in both cases, from ARMv8.1's LSE (Large System Extensions). The stronger ordering adds only the al suffix, meaning acquire and release |
The opcodes pick the ordering, and the hardware does the rest. We've now seen what release and acquire cost on ARM, which was built for weak ordering. x86 behaves very differently, and that difference is why the bug from section 3 can pass every test you own.
06Why x86 hides the bug
6.1Different processors reorder differently
The model is the same everywhere, but processors differ in how much they reorder before the programmer asks for anything. x86-64 follows a set of rules called total store order (TSO). Each core's store buffer drains in the order the stores went in, so a store can't overtake an earlier store, and in our example B can't land before A. The one reordering TSO allows is a later load passing an earlier store to a different address, because the load doesn't wait for the store sitting in the buffer. ARM promises far less:
| Architecture | Reorders on its own | Consequence |
|---|---|---|
| x86-64 (TSO) | A load can pass an earlier store to a different address, and nothing else | Most under-synchronised code accidentally works |
| ARMv8 / AArch64 | Almost anything: stores with stores, loads with loads, loads with stores | The same code fails, months later |
| POWER, RISC-V | Also weak, so ordering bugs show up there too | Rare in production, brutal when hit |
Code with the release and acquire left out can therefore pass every test on an Intel laptop and fail on an AWS Graviton instance, an M-series Mac or an Ampere server:
// WRONG — and it passes everything on x86-64
data = compute(); // plain store
ready.store(true, std::memory_order_relaxed);?Why does this work on x86?
Because TSO won't reorder those two stores, an x86 reader that observes ready also observes data. (The compiler may still swap them, so even x86 gives no guarantee.) ARM reorders them happily. It's the same source and the same compiler, and the bug waits for different silicon. Running the program from section 4.3 on an x86-64 machine is the quickest way to see the difference: the relaxed line should stay at 0 there for the same reason.
The classic casualty is double-checked locking, a way to create a lazy singleton, an object built the first time somebody asks for it. Each caller checks a shared pointer without taking a lock. If it's null, the caller takes the lock, checks again, builds the object and stores the pointer. The idea was widely published with a plain pointer, which is broken on any machine that reorders: a caller that skips the lock can see a non-null pointer to a half-constructed object, and that's the publication bug again. Java 1.5 fixed it by changing the memory model so that a volatile field works. C++ needs acquire and release, or std::call_once.
6.2volatile isn't part of this model
You might wonder whether volatile, which we used in section 1, can do this job. It tells the compiler not to remove or cache an access. It says nothing about the processor, nothing about atomicity, and nothing about order between different variables, and it creates no happens-before edge. The plain counter in section 1 was volatile and still lost hundreds of thousands of increments.
Java volatile | C/C++ volatile | C++ std::atomic | |
|---|---|---|---|
| Stops the compiler caching the access | Yes | Yes | Yes |
| Orders other memory accesses | Roughly sequential consistency | No | Yes, per the ordering you choose |
| Meant for | Thread synchronisation | Device registers, signal handlers | Thread synchronisation |
Java's meaning is where the confusion comes from.
That leaves the objection every performance-minded programmer raises: if seq_cst is the strongest ordering and the safest default, surely it's the slow one.
07What ordering costs
7.1What the strongest ordering costs
Single loads and stores are hard to time honestly. At 0.2 to 0.4 ns each they're at the edge of what a timing loop can resolve, and a compiler can lift a plain load out of the loop altogether, which makes the number meaningless. For them, the assembly in section 5.2 is better evidence: a relaxed load is ldr and an acquire load is ldapr, one instruction each. A read-modify-write takes longer and can be timed reliably, one fetch_add at a time on a single thread:
Sequential consistency was free here, with no contention. The two versions compile to a single ARMv8.1 atomic each, and the stronger ordering is a different opcode and not extra work, as the ldadd and ldaddal lines in section 5.2 showed.
On older ARM without LSE, or on x86 where a seq_cst store needs an xchg or mfence instruction, there is a real gap, and you should measure it on your own hardware. Set it beside the cost in the next subsection before deciding it matters.
7.2Contention costs far more
Contention means several threads fighting over the same variable. The cache line holding it has to travel between their cores on every update. Chapter 14 measures the comparison that matters. It uses its own test program, so its uncontended number differs from the one above:
| Operation | Cost |
|---|---|
fetch_add, uncontended | 3.0 ns |
fetch_add, eight threads | 18.4 ns |
| Compare-exchange loop, under the same contention | 162 ns |
A compare-exchange loop (CAS) is the other way to update an atomic: write the new value only if the variable still holds the value you last read, and go round again if another thread got there first.
Put the numbers to work on one million increments:
| Uncontended fetch_add | 1,000,000 × 3.0 ns | 3 ms |
| fetch_add, eight threads | 1,000,000 × 18.4 ns | 18.4 ms |
| Compare-exchange loop, eight threads | 1,000,000 × 162 ns | 162 ms |
| Gap between relaxed and seq_cst, single thread | 1,000,000 × 0.004 ns | 0.004 ms |
| contention against ordering | 6×–54× vs ~0% | |
Eight threads sharing one counter make each fetch_add about six times slower than an uncontended one (18.4 ÷ 3.0), and a compare-exchange loop under the same contention costs about fifty times as much (162 ÷ 3.0 is 54). Against that, the choice of ordering made no measurable difference. So when atomics are slow, look first at how many threads are hitting the same variable, and only then at the orderings.
We have enough now to turn this into rules, and into ways to find the bugs we can't reason our way out of.
08Using it in practice
8.1Finding the bugs
Each question the chapter raised has a tool that answers it on your own code.
# Which processor is this? arm64 or aarch64 is ARM; x86_64 is x86. (section 6)
uname -m
# Is there a data race anywhere in this program? (section 3)
clang++ -fsanitize=thread -O1 -g app.cpp # ThreadSanitizer: finds what review doesn't
# What did each ordering compile to on this target? (section 5)
clang++ -O2 -S -o - app.cpp | grep -E 'ldar|ldapr|stlr|dmb|lock|mfence|xchg'8.2Rules that hold up
- Default to
seq_cst. It's the C++ default because reasoning about anything weaker is hard, and section 7 found no measurable cost for it on a modern ARM core without contention. - Use relaxed for counters whose value never gates control flow. Metrics, progress indicators, statistics.
- Use acquire/release in pairs, and write the pairing in a comment. An unpaired acquire is a bug that compiles.
- Never use
volatilefor thread synchronisation. - Test on ARM.
8.3What you trade for what
| You choose | You get | You pay | When the bill arrives |
|---|---|---|---|
seq_cst (the default) | One order every thread agrees on, and the easiest code to reason about | Possibly a stronger instruction, such as ldar or mfence, on some targets | As a small cost you can measure, but only if you look |
acquire / release pairs | The publication pattern from section 4 at the lowest cost that works | Both halves must be correct, and nothing warns you if one is missing | As a stale read on ARM, months later |
relaxed | Atomicity with no ordering | Nothing else is ordered with it | As a half-built object if the value gates anything |
| Plain variables | The fastest code | A data race and undefined behaviour | As a compiler upgrade that breaks the program |
8.4Symptom, cause, fix
| Symptom | Likely cause | Fix |
|---|---|---|
| Passes on x86, fails on ARM | Missing acquire/release; TSO was hiding it | Release store, acquire load; run tests on ARM |
| Reader sees a non-null pointer to a half-built object | Publication with a relaxed or plain store | Release the pointer, acquire it on the reader |
| Racy code broke after a compiler upgrade | A data race is undefined behaviour, and the optimiser relied on that | Make the shared variable atomic |
A volatile flag between threads misbehaves | volatile orders nothing | std::atomic |
| Atomics are slow under load | Contention, not ordering | Reduce sharing before weakening ordering |
09Summary
- A plain shared counter loses updates, because
plain + 1is three steps and two threads can interleave them. An atomic makes it one. - Neither the compiler nor the CPU runs your code in order. They only promise that one thread can't tell.
- Store buffers let later stores become visible first. Your own thread reads through the buffer, and other threads don't.
- A reader can see the flag and still read the old data. In the publication example, B reached the cache ahead of A.
- A data race is undefined behaviour in C++. The optimiser may assume it never happens, so a racy program has no defined meaning at all, which is worse than reading a stale value.
- Happens-before is the whole model. It comes from program order plus synchronises-with edges, and atomics, mutexes and thread start and join create those edges.
- Acquire and release only work in pairs, on the same variable. An unpaired acquire compiles and guarantees nothing.
- Publication is the pattern to know. Construct, release-store the pointer, acquire-load it on the reader. Run two million rounds and the relaxed version fails while the release and acquire version does not.
- On ARMv8, orderings are opcodes, not barriers. A
seq_cststore is the samestlras a release store, and aseq_cstload isldar. - x86 hides ordering bugs that ARM exposes. TSO keeps stores in order, so test on ARM.
volatileisn't synchronisation in C or C++. - Ordering is cheap and contention isn't.
fetch_addcost 1.596 ns relaxed and 1.592 nsseq_cst, and 18.4 ns across eight threads.
10Build this
Reproduce store-buffer reordering yourself. Two threads, the Dekker shape, where each thread stores to one variable and then loads the other:
thread 1: x = 1; r1 = y;
thread 2: y = 1; r2 = x;Sequential consistency forbids r1 == 0 && r2 == 0: one of the two stores must come first, so the other thread has to see it.
- Run it a few million rounds with relaxed atomics, using the layout from the section 4.3 program (a fresh
xandyper round on separate cache lines), and count how often that impossible result appears. Any machine with store buffers shows it, x86 included, because a load passing an earlier store is exactly what TSO allows. Expect a large count. - Then switch to
seq_cstand watch the count go to zero. - Compile both with
-Sand find the instructions that changed.
Seeing an impossible outcome happen on hardware you own is what makes this stop being abstract.
11Interview questions
beginnerWhat is a data race, and why is it worse than a stale read?›
Two threads accessing one location, at least one writing, with no synchronisation between them. In C++ it's undefined behaviour: the compiler may assume it can't happen and optimise on that basis.
So the outcome isn't "you read an old value". Surrounding code may be transformed in ways that make no sense. That's why racy programs work for years and then break on a compiler upgrade.
intermediateWhat do acquire and release guarantee?›
Release on a store means nothing before it in program order moves after it. Acquire on a load means nothing after it moves before it. Pair them on the same variable and everything the releasing thread wrote before the store is visible to the acquiring thread after the load.
They only work in pairs. An acquire load with no matching release store compiles fine and guarantees nothing at all.
intermediateIs volatile enough for sharing a flag between threads in C++?›
No. It stops the compiler eliding or caching the access and does nothing else: no processor ordering, no atomicity for wide types, no happens-before edge.
Java's volatile does imply ordering, which is where the confusion comes from. In C++ use std::atomic. volatile is for device registers and signal handlers.
deepYour code passes on x86 and fails on ARM. What's the likely class of bug?›
Missing acquire/release. x86-64 is total-store-ordered, so it only reorders a store followed by a load to a different address. A plain store followed by a flag store stays in order, and most under-synchronised code accidentally works. ARM reorders far more freely.
Concretely, data = compute(); ready.store(true, relaxed); is fine on x86 and broken on ARM, because a reader can observe ready before data. A two-million-round test of that shape shows thousands of stale reads on ARM and none with a release and acquire pair. Use a release store and an acquire load, and run your test suite on ARM routinely.
deepHow much does seq_cst cost compared to relaxed?›
Less than people assume, and you should measure instead of guessing. A single-threaded fetch_add takes about 1.596 ns relaxed and 1.592 ns seq_cst, indistinguishable, because both compile to a single ARMv8.1 LSE atomic with the ordering encoded in the opcode (ldadd against ldaddal). The compiler's assembly output shows a seq_cst store emitting exactly the same stlr as a release store.
Contention is the expensive thing. The same operation reaches 18.4 ns across eight threads, and a compare-exchange loop 162 ns. Reduce sharing before weakening ordering.
12Go deeper
On ARMv8, what does a seq_cst store compile to?›
stlr, the same instruction as a release store. The stronger ordering needs
no extra barrier on the store side. A seq_cst load is the one that differs
from an acquire load: it uses ldar where acquire uses ldapr.
You see an acquire load with no release store anywhere. Bug?›
Yes. They only mean something as a pair on the same variable. An unpaired acquire compiles and guarantees nothing.
Why can code with a data race start failing after a compiler upgrade, with no source change?›
Because a data race is undefined behaviour, not a defined stale read. The optimiser may assume no race exists and transform surrounding code on that basis, and a new compiler may transform it differently.
Cheapest way to find ordering bugs in a codebase tested only on x86?›
Run the test suite on ARM. A Graviton instance or an M-series laptop exposes a class of bug that x86 structurally cannot.
Two long talks, still the clearest spoken explanation of the C++ memory model.
Jeff Preshing's series, with a runnable experiment for every claim. Start with "Memory Barriers Are Like Source Control".
Exhaustive and practical, written by people who had to make one model work across every architecture Linux supports. Read it once for the ASCII diagrams of store buffers misbehaving.
The precise wording. Dry, and the authority when an argument needs settling.
Chapter 26 shows the same lost update on a shared counter, step by step in assembly. Chapter 28 covers the hardware atomic instructions (test-and-set, compare-and-swap, fetch-and-add) that locks are built from. Free online at ostep.org.
13Related chapters
Where the 83 ns fetch from main memory comes from, and how coherence moves lines between cores. Chapter 02.
What a mutex builds on top of these orderings. Chapter 13.
Where the contention numbers in section 7 come from, and what you build once ordering is settled. Chapter 14.