You write a four-line C program called hello.c that calls puts("hello"), compile it with gcc -O2 -o hello hello.c, type ./hello and press Enter. "hello" appears and the program exits, and the whole trip from Enter to exit takes about 150 microseconds. It's natural to picture the compiled file as the entire program: a block of machine code that the computer copies into memory and starts running.
That picture leaves out puts. Your file contains a call to it, but the code that writes to the screen lives in the C library, a separate file on disk that gcc never copied in. When you press Enter, something has to find that file, bring it into memory and tell your program where puts ended up, all before your first line runs. If any of that fails, the program dies before it begins, and the error message often points at the wrong thing.
A program called the dynamic loader, ld.so, does this work, and the kernel starts it before your own code gets a turn. This chapter follows one run of hello from Enter to the first line of main, and for each step asks what it does, why it has to exist, what it costs and how it fails. We start by asking a real program what it needs, work out why that need exists, read the file that states it, and then watch the whole sequence happen.
01What a program needs before it runs
1.1Asking ls what it needs
Let's begin with something real. ls is small and has run on your machine thousands of times, so it makes a fair test subject. These commands run inside a throwaway Ubuntu 24.04 Docker container so that everyone sees the same files, but any Linux machine gives the same kind of answer.
The container runs three commands, and each needs a word of explanation.
The first, ldd /bin/ls, lists the shared libraries that ls needs. A shared library is a file of ready-made functions that many programs use. The most important one is libc.so.6, the C library, which holds puts, malloc, read and thousands of other functions.
The second runs ls with LD_DEBUG=libs in front of it. That's an environment variable, a named setting that a program inherits from the shell that started it. Writing NAME=value command sets it for that one command, and this particular one asks the loader to describe its search for each library. The loader writes that description to standard error, so 2>&1 >/dev/null sends standard error down the pipe and throws away ls's normal output, and head -12 keeps the first dozen lines.
The third sets LD_PRELOAD, another environment variable, which names a library to load before all the others. We give it a file that doesn't exist, to see what the loader does about it.
docker run --rm ubuntu:24.04 sh -c '
ldd /bin/ls
echo
LD_DEBUG=libs /bin/ls / 2>&1 >/dev/null | head -12
echo
LD_PRELOAD=/nope.so /bin/ls / 2>&1 | head -3' linux-vdso.so.1 (0x0000ffff9aa17000)
libselinux.so.1 => /lib/aarch64-linux-gnu/libselinux.so.1 (0x0000ffff9a940000)
libc.so.6 => /lib/aarch64-linux-gnu/libc.so.6 (0x0000ffff9a780000)
/lib/ld-linux-aarch64.so.1 (0x0000ffff9a9da000)
libpcre2-8.so.0 => /lib/aarch64-linux-gnu/libpcre2-8.so.0 (0x0000ffff9a6d0000)
14: find library=libselinux.so.1 [0]; searching
14: search cache=/etc/ld.so.cache
14: trying file=/lib/aarch64-linux-gnu/libselinux.so.1
14:
14: find library=libc.so.6 [0]; searching
14: search cache=/etc/ld.so.cache
14: trying file=/lib/aarch64-linux-gnu/libc.so.6
14:
14: find library=libpcre2-8.so.0 [0]; searching
14: search cache=/etc/ld.so.cache
14: trying file=/lib/aarch64-linux-gnu/libpcre2-8.so.0
14:
ERROR: ld.so: object '/nope.so' from LD_PRELOAD cannot be preloaded (cannot open shared object file): ignored.
bin
bootA tiny program like ls depends on three libraries: libselinux, libc and libpcre2. The list also holds the loader itself, /lib/ld-linux-aarch64.so.1, and linux-vdso.so.1, which isn't a file on disk at all. It's a small piece of code the kernel places in every process's memory so that a few system calls can run without entering the kernel (chapter 07 calls it the vDSO). Each line ends with an address, the place in memory where that piece landed for this run. Those numbers differ on every run, and section 2 explains why.
1.2What the second and third command show
In the second block the loader narrates its search. For each library it first checks a cache file, /etc/ld.so.cache, a prebuilt index from library names to file paths that saves it from searching directories one by one. Then it tries the path the cache gave it. The 14 at the start of each line is the process ID, the number the kernel gave this run of ls.
The third block shows LD_PRELOAD at work. It's a hook into the loader: any library you name there gets loaded before all the others, and its functions win over the real ones. A file that doesn't exist only produces a warning, and ls goes on to list / without it. That hook is how profilers and replacement allocators such as jemalloc get into a program without rebuilding it, and section 6 shows what it can do.
Everything that follows explains what happened between you pressing Enter and ls printing its first line: how the kernel starts the loader, what's in the file it reads, and what it costs every time any process starts.
1.3The program we'll follow
ls already needs three libraries, which is more than we want to draw. For the rest of the chapter we'll follow hello, the smallest program that still has the problem. Here is the source, and what ldd says about the compiled file:
// hello.c
#include <stdio.h>
int main(void) { puts("hello"); return 0; }$ gcc -O2 -o hello hello.c
$ ./hello
hello
$ ldd ./hello
linux-vdso.so.1 (0x0000ffff96b09000)
libc.so.6 => /lib/aarch64-linux-gnu/libc.so.6 (0x0000ffff968e0000)
/lib/ld-linux-aarch64.so.1 (0x0000ffff96acc000)hello needs only libc, plus the loader and the vDSO page that every process gets. The #include <stdio.h> line only declares puts, telling the compiler what its arguments look like. It doesn't contain its code. So the compiled hello holds a call to a function whose code is in another file, and the obvious question is where the address for that call comes from.
02Why the program doesn't contain puts
2.1A call to an address nobody knows
A recipe card says "add the tomato sauce from the pantry" and "use the stock on shelf two". It doesn't contain the sauce or the stock. Before anyone cooks, an assistant reads the card, walks to the pantry, fetches every item it names and lays them on the counter. Your program is the recipe card, and the loader is the assistant.
Here is how the card gets written. gcc turns your call to puts into a branch instruction, bl puts on the ARM processors these examples run on, which means "jump to the address of puts and come back afterwards". puts is a symbol, a name for a function or variable that one piece of code defines and another uses. But the compiler only ever saw hello.c. It never saw the C library's code, so it can't write down the address of puts. Somebody has to supply that number, and there are two ways to do it: at build time, or at run time.
2.2Static linking: copy the code in
Doing it at build time is the simplest answer. A tool called the linker (ld, which gcc runs for you at the end of a build) combines the compiled pieces into one executable. It can copy the machine code for puts, and everything puts calls, into your executable and patch the branch to point at the copy. This is static linking, and it works. If you build hello with gcc -static, the result needs no libraries: ldd reports "not a dynamic executable".

It also makes the file much bigger. Static hello is 639,808 bytes, against 70,312 for the dynamic one.
?Why not do that for every program?
Think about what it means across a whole system. Every program on the disk carries its own copy of the C library, and every running process holds its own copy in RAM. And when a security fix lands in the library, every program has to be relinked to pick it up.
2.3Dynamic linking: one copy, many processes
So most systems share instead. One copy of libc.so.6 sits on disk, and each process that needs it asks the kernel to make the file's bytes appear at some range of the process's memory addresses. This is called mapping the file. The kernel manages memory in pages, chunks of 4 KB on most systems, and when many processes map the same file, it keeps one copy of each page in physical memory and shows that copy to all of them. A thousand processes then cost roughly one copy of libc. Sharing one library file this way, and connecting each program to it when the program starts, is dynamic linking, and it creates a new problem.
hello can't have the address of puts written into it, because libc lands somewhere different in each process. That's partly on purpose. Linux uses ASLR (address space layout randomisation): it places code at random addresses on every run, so an attacker who has found a bug can't rely on knowing where any function lives. So nobody knows the address of puts until the process starts.
The fix is a level of indirection. hello's call goes through an entry in a table inside the program's own memory (we'll meet it as the GOT in section 3), and the table starts out empty. Before main runs, someone has to find each library by name, map it, work out where each function landed, and write those addresses into the table. That someone is the dynamic loader, ld.so, which the kernel starts before your program gets control. Each fill-in is a relocation, an instruction of the form "write this address in this spot". Starting hello takes 1,214 of them, by the loader's own count. Only a handful belong to hello itself, and almost all the rest are libc fixing up pointers to its own code for wherever it landed.
| Static linking | Dynamic linking | |
|---|---|---|
| When addresses are filled in | At build time, by the linker | At process start (or first call), by ld.so |
| Copies of libc | One per program, on disk and in RAM | One file, shared pages |
| A libc security fix | Relink every program | Replace one file |
| Who runs first | Your _start | ld.so |
The loader can't guess any of this. Each file has to carry its own instructions, so the next question is how a file can say "I need libc, and here are the spots to fix".
03Inside an ELF file
Linux stores executables and libraries in a format called ELF, the Executable and Linkable Format. An ELF file is a header, two tables that describe the rest of the file, and the data those tables point at. We'll open hello and read what the loader reads.
3.1The header, and the 64 bytes the kernel reads first
Every ELF file opens with a fixed header, 64 bytes long on 64-bit targets. The kernel checks the magic number, the first four bytes, 7f 45 4c 46, which spell \x7fELF. Then it checks the file's type and the processor it's for. The header also holds two offsets, positions in the file where the two tables start: e_phoff for the program headers and e_shoff for the section headers. Section 3.2 explains why there are two.

Those numbers are from a hello built with plain gcc -O2 on Ubuntu 24.04. Its type field says DYN, the type used for shared libraries, where you might expect EXEC, the type for a program that must sit at one fixed address. That's because a default Ubuntu build is a position-independent executable (PIE): its code works wherever in memory it's placed. A PIE lets the kernel put the program itself at a random address, the same ASLR it applies to libraries, and Linux treats it almost exactly like a shared library. e_entry, 0x680 here, is where execution begins, counted from wherever the file ends up being placed.
3.2Program headers are for loading, section headers are for tools
An ELF file describes itself twice, because two different readers want different things. The kernel and the loader only need to know what to put in memory, where, and with which permissions. A linker or debugger wants to know what each stretch of bytes means.
| Table | Describes | Read by |
|---|---|---|
| Program headers (segments) | What to map, where, and with which permissions | The kernel and ld.so |
| Section headers | What the bytes mean: .text (the code), .dynsym (the dynamic symbols), .rela.plt (the function-call relocations) | The linker, objdump, gdb |

A segment is a stretch of the file that gets mapped as one piece. Here are hello's segments:
$ readelf -lW hello # trimmed
Type Offset VirtAddr FileSiz MemSiz Flg Align
PHDR 0x000040 0x0000000000000040 0x0001f8 0x0001f8 R 0x8
INTERP 0x000238 0x0000000000000238 0x00001b 0x00001b R 0x1
[Requesting program interpreter: /lib/ld-linux-aarch64.so.1]
LOAD 0x000000 0x0000000000000000 0x0008b4 0x0008b4 R E 0x10000
LOAD 0x00fd90 0x000000000001fd90 0x000280 0x000288 RW 0x10000
DYNAMIC 0x00fda0 0x000000000001fda0 0x0001f0 0x0001f0 RW 0x8
GNU_STACK 0x000000 0x0000000000000000 0x000000 0x000000 RW 0x10
GNU_RELRO 0x00fd90 0x000000000001fd90 0x000270 0x000270 R 0x1There are two LOAD segments, which are the ones the loader maps. One is marked R E (readable and executable: the code) and the other RW (readable and writable: the data). No segment is both writable and executable, so data can't be run as code. INTERP names the program that runs first, which will matter a great deal in section 4. DYNAMIC points at the loader's instructions, covered next. GNU_RELRO marks the part of the writable segment that goes read-only once relocation is done, and section 5 explains why that's worth doing.
?Why is this file 70 KB for 2 KB of code?
Because of that 0x10000 in the Align column. A segment can only be mapped starting at a page boundary, and ARM Linux kernels can be built with 4 KB, 16 KB or 64 KB pages. The aarch64 toolchain aligns segments to 64 KB so the same binary works on all three, and the file is padded to match.
3.3.dynamic, relocations, and the GOT
The ELF specification names segment types with a PT_ prefix, which readelf leaves off, so the DYNAMIC line above is PT_DYNAMIC and INTERP is PT_INTERP. PT_DYNAMIC points at the .dynamic section, an array of tag and value pairs that works as the loader's table of contents. Hello's has 27 entries. These are the ones that matter:
| Tag | What it tells the loader |
|---|---|
NEEDED libc.so.6 | The dependency list, by soname, a library's name with a version number in it so an incompatible new version gets a different name |
SYMTAB, STRTAB, GNU_HASH | The dynamic symbol table, its strings, and the hash table for looking names up |
RELA, JMPREL | Two relocation tables: .rela.dyn for data, .rela.plt for function calls |
PLTGOT | The address of the GOT, the Global Offset Table: the table of addresses that hello's calls go through |
FLAGS BIND_NOW, FLAGS_1 NOW PIE | Bind everything at startup. Section 5 comes back to this. |
A relocation is an instruction to the loader: at this offset, write the address of that symbol, plus this addend. Here's readelf -r on a lazily linked build of the same hello (section 5 explains "lazily"):
Relocation section '.rela.dyn' at offset 0x480 contains 8 entries:
Offset Type Symbol's Name + Addend
000000000001fdc8 R_AARCH64_RELATIVE 790
000000000001ffc8 R_AARCH64_GLOB_DAT __cxa_finalize@GLIBC_2.17 + 0
...
Relocation section '.rela.plt' at offset 0x540 contains 5 entries:
0000000000020000 R_AARCH64_JUMP_SLOT __libc_start_main@GLIBC_2.34 + 0
0000000000020020 R_AARCH64_JUMP_SLOT puts@GLIBC_2.17 + 0That's three kinds of relocation, with three costs:
| Type | What the loader does | Cost |
|---|---|---|
RELATIVE | Adds the load base and stores | No lookup at all |
GLOB_DAT | Looks up a data or function symbol | A hash, a walk over every loaded object, a string compare |
JUMP_SLOT | Looks up a function and fills its GOT slot | Same lookup, but it can be put off until the first call |
RELATIVE relocations are pointers within the same file, such as "this variable holds the address of that function", so the loader only has to add the address where the file landed. The other two name a symbol, and finding the symbol is where the time goes.
3.4GNU hash: a Bloom filter in front of every lookup
A symbol lookup asks each loaded object in turn, "do you define puts?" (A loaded object is the executable or any library the loader has mapped.) Most of them say no, so making "no" cheap is what matters. DT_GNU_HASH does this with a Bloom filter, a small array of bits that can say "definitely not here" after testing just a couple of bits, and sometimes says "maybe", in which case the real table has to be checked.

Here is one lookup of puts against hello's executable and then libc:
puts. It hashes the name once, then asks each object in search order: the executable first, then libc.Here is the same lookup in glibc 2.39's source. Two details the picture leaves out: the two bits come from different parts of the hash (hashbit1 from its low bits, hashbit2 after shifting it right), and the lowest bit of each chain entry marks the end of the chain.
const ElfW(Addr) *bitmask = map->l_gnu_bitmask;
if (__glibc_likely (bitmask != NULL))
{
ElfW(Addr) bitmask_word
= bitmask[(new_hash / __ELF_NATIVE_CLASS)
& map->l_gnu_bitmask_idxbits];
unsigned int hashbit1 = new_hash & (__ELF_NATIVE_CLASS - 1);
unsigned int hashbit2 = ((new_hash >> map->l_gnu_shift)
& (__ELF_NATIVE_CLASS - 1));
/* Two bits must both be set, or this object can't define the name. */
if (__glibc_unlikely ((bitmask_word >> hashbit1)
& (bitmask_word >> hashbit2) & 1))
{
Elf32_Word bucket = map->l_gnu_buckets[new_hash % map->l_nbuckets];
if (bucket != 0)
{
const Elf32_Word *hasharr = &map->l_gnu_chain_zero[bucket];
do
/* Compare 31 bits of hash before any strcmp. */
if (((*hasharr ^ new_hash) >> 1) == 0)
{ /* ... check_match(): the actual string compare ... */ }
while ((*hasharr++ & 1u) == 0); /* low bit ends the chain */
}
}Note the __glibc_unlikely on the Bloom test. glibc's authors expect most objects to answer "no", so the compiler is told to lay out the code with that case as the fast path. In a C++ program with a dozen libraries, most probes probably end at one 64-bit load and two shifts. For hello's two objects the filter barely matters, and it pays off as the number of loaded objects grows.
That completes what the file says: which libraries it needs, which spots to fix, and how to find names quickly. What's left is to watch who reads all this, and in what order, when you press Enter.
04From execve to main
When you press Enter, the shell asks the kernel to start hello with the system call execve (the shell has first made a copy of itself with fork, and the copy calls execve). That call replaces the copy's program with the one in the file. Everything between the kernel taking over and your first line of main is the path we want to understand. Here is the whole thing for a dynamically linked hello, with the three players in it: the kernel, the loader and your code.
execve("./hello"). The kernel opens the file and reads its program headers. One of them, INTERP, names another program that must run first: the loader.Next we take the frames one at a time, starting with the kernel's part. One thing to notice already is that hello and libc were never "loaded" by copying. Mapping makes their bytes appear in the process, and the pages are brought in from disk only when something touches them.
4.1The kernel: load_elf_binary
execve walks a list of binary formats until one claims the file. For ELF that's load_elf_binary in fs/binfmt_elf.c. It reads the program headers, and the first thing it hunts for is PT_INTERP:
for (i = 0; i < elf_ex->e_phnum; i++, elf_ppnt++) {
char *elf_interpreter;
...
if (elf_ppnt->p_type != PT_INTERP)
continue;
retval = -ENOEXEC;
if (elf_ppnt->p_filesz > PATH_MAX || elf_ppnt->p_filesz < 2)
goto out_free_ph;
retval = -ENOMEM;
elf_interpreter = kmalloc(elf_ppnt->p_filesz, GFP_KERNEL);
...
retval = elf_read(bprm->file, elf_interpreter, elf_ppnt->p_filesz,
elf_ppnt->p_offset);
...
interpreter = open_exec(elf_interpreter); /* the path from .interp */
kfree(elf_interpreter);
retval = PTR_ERR(interpreter);
if (IS_ERR(interpreter))
goto out_free_ph; /* ENOENT comes from here */So the path inside .interp gets opened by the kernel, from the kernel's view of the filesystem. Keep that last comment in mind for section 6.1.
Next it maps each PT_LOAD segment at a random base address (a PIE has type ET_DYN, the DYN from section 3.1, so it gets ASLR), maps the interpreter the same way, and picks where to jump. In the code below, load_bias is that random base:
e_entry = elf_ex->e_entry + load_bias; /* your program's entry */
...
if (interpreter) {
elf_entry = load_elf_interp(interp_elf_ex, interpreter,
load_bias, interp_elf_phdata, &arch_state);
if (!IS_ERR_VALUE(elf_entry)) {
interp_load_addr = elf_entry;
elf_entry += interp_elf_ex->e_entry; /* ...but we start in ld.so */
}
...
} else {
elf_entry = e_entry; /* static: straight to _start */
}
...
retval = create_elf_tables(bprm, elf_ex, interp_load_addr,
e_entry, phdr_addr);
...
START_THREAD(elf_ex, regs, elf_entry, bprm->p);So there are two entry points. The CPU starts in ld.so, and your program's own entry goes on the stack as AT_ENTRY for ld.so to jump to later.
4.2The auxiliary vector: the kernel's notes
create_elf_tables writes the auxiliary vector onto the stack, next to argv and envp. It's a list of facts the kernel already knows and the loader would otherwise have to work out, such as where the vDSO is, how big a page is, and where hello's program headers ended up. glibc will print it for you:
$ LD_SHOW_AUXV=1 ./hello # trimmed
AT_SYSINFO_EHDR: 0xffffbc9dc000 # the vDSO from chapter 07
AT_HWCAP: efb3ffff # CPU features (section 5.3)
AT_PAGESZ: 4096
AT_PHDR: 0xaaaac8e80040 # where your program headers ended up
AT_BASE: 0xffffbc99f000 # where ld.so was mapped
AT_ENTRY: 0xaaaac8e80680 # 0x680 from the header, plus the bias
AT_RANDOM: 0xffffe499fd78 # 16 bytes: stack canary, pointer guardAT_HWCAP is a bitmask of CPU features, which section 5.3 uses. AT_RANDOM points at 16 random bytes that libc turns into secret values: a "canary" placed on the stack so a buffer overflow that overwrites it gets caught, and a key for scrambling saved function pointers. AT_ENTRY is the 0x680 from the ELF header plus the random base, which is hello's _start. ld.so doesn't have to parse your ELF file again. Linux already did, and this is it handing over its notes.
4.3ld.so: map the libraries, apply the relocations
Now the loader is running. Here's the complete strace of the dynamic hello, with only arguments trimmed. (strace prints every system call a program makes. openat opens a file, mmap maps a file or a range of memory into the process, and mprotect changes the permissions on a range that's already mapped.)
execve("./hello", ["./hello"], 0xffffc129ca80 /* 6 vars */) = 0
brk(NULL) = 0xaaaadd59a000
mmap(NULL, 8192, PROT_READ|PROT_WRITE, ...)
faccessat(AT_FDCWD, "/etc/ld.so.preload", R_OK) = -1 ENOENT
openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
fstat(3, ...); mmap(NULL, 10999, PROT_READ, MAP_PRIVATE, 3, 0); close(3)
openat(AT_FDCWD, "/lib/aarch64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
read(3, "\177ELF\2\1\1\3\0\0\0\0\0\0\0\0\3\0\267\0..., 832) = 832
mmap(NULL, 1892240, PROT_NONE, ...) # reserve the whole span first
mmap(0xffffb6ca0000, 1826704, PROT_READ|PROT_EXEC, MAP_FIXED, 3, 0)
munmap(...); munmap(...); mprotect(..., PROT_NONE)
mmap(0xffffb6e4d000, 20480, PROT_READ|PROT_WRITE, MAP_FIXED, 3, 0x19d000)
mmap(0xffffb6e52000, 49040, PROT_READ|PROT_WRITE, MAP_FIXED|MAP_ANONYMOUS)
close(3)
set_tid_address(...); set_robust_list(...); rseq(...)
mprotect(0xffffb6e4d000, 12288, PROT_READ) # libc's RELRO
mprotect(0xaaaad2a2f000, 4096, PROT_READ) # hello's RELRO
mprotect(0xffffb6eaa000, 8192, PROT_READ) # ld.so's RELRO
...
write(1, "hello\n", 6) = 6
exit_group(0) = ?Read it in order. After execve, the loader looks for /etc/ld.so.preload and opens /etc/ld.so.cache, its first lookups. It then opens libc.so.6, reads the first 832 bytes (the ELF header and program headers) and makes a run of mmap calls, one per segment. The three mprotect(..., PROT_READ) calls near the end are full RELRO being applied to libc, to hello and to ld.so itself, which section 5 explains. The write is hello's own output, and exit_group ends the process.
That's 32 syscalls in total. A static build makes 15.
?Why does nobody open hello or ld.so?
Because the kernel mapped both already, in load_elf_binary. The loader opens only what the kernel didn't: the cache and the libraries.
Two more things stand out:
- The loader checks
/etc/ld.so.preloadbefore it reads the cache. It's a system-wideLD_PRELOADthat persists across reboots, and a favourite of rootkits, malware that hides itself by replacing the functions tools likelsuse. - It reserves libc's whole address range as
PROT_NONEfirst, then maps the segments into it withMAP_FIXED, so nothing else can land in the gap between them.
Between the last close(3) and the first mprotect is where relocations get applied, with no syscalls to see. LD_DEBUG=statistics counts them:
$ LD_DEBUG=statistics ./hello
3992: number of relocations: 98
3992: number of relocations from cache: 7
3992: number of relative relocations: 1116These are the 1,214 from section 2.3. Almost all of the 1,116 relative ones are libc fixing up its own pointers, and the 98 symbol relocations are the ones that cost lookups. On some architectures, x86-64 among them, glibc also prints how many CPU cycles startup took; the aarch64 build of glibc prints no such line.
4.4Constructors, __libc_start_main, and finally main
After relocation ld.so runs each library's DT_INIT_ARRAY, an array of constructors: functions a library marks to run once at start-up, before any of its code is called. libc's run first, because the executable depends on it. Then the loader jumps to AT_ENTRY. That's your _start: a few instructions from crt1.o (a small piece of startup code gcc links into every program) that call __libc_start_main. The strings LD_DEBUG prints in section 5.2 are right here in the source:
#ifdef SHARED
if (__builtin_expect (GLRO(dl_debug_mask) & DL_DEBUG_IMPCALLS, 0))
GLRO(dl_debug_printf) ("\ninitialize program: %s\n\n", argv[0]);
if (init != NULL)
/* This is a legacy program which supplied its own init routine. */
(*init) (argc, argv, __environ MAIN_AUXVEC_PARAM);
else
/* This is a current program. Use the dynamic segment to find
constructors. */
call_init (argc, argv, __environ); /* your DT_INIT_ARRAY */
...
if (__glibc_unlikely (GLRO(dl_debug_mask) & DL_DEBUG_IMPCALLS))
GLRO(dl_debug_printf) ("\ntransferring control: %s\n\n", argv[0]);
#endif
__libc_start_call_main (main, argc, argv MAIN_AUXVEC_PARAM);
}__libc_start_call_main calls main and passes its return value to exit. That init != NULL branch is also behind one of the most common deployment errors on Linux (section 6.2).
Hello made 98 symbol lookups before its first line ran, and a big C++ service can make tens of thousands. Many of those are for functions the program may never call. Looking all of them up in advance isn't required, and that's the next idea.
05Lazy binding and IFUNC
Relocations for function calls don't have to be done up front. With lazy binding, each one waits until the function is first called. And some symbols aren't looked up at all: for those, the loader runs a function to decide what the answer is.
5.1The PLT stub and its GOT slot
Calls to puts go through a small stub in the PLT (procedure linkage table), which has one stub per imported function. Each stub jumps through a GOT slot, and that slot initially points back into the loader. The first call resolves puts and patches the slot, and every call after that goes straight to libc.
?Is our hello lazily bound?
No, because of the FLAGS BIND_NOW in section 3.3. Ubuntu 24.04's gcc links with -z now by default, so every JUMP_SLOT is resolved at startup and the GOT becomes read-only. Making the GOT read-only after startup is called full RELRO (relocation read-only). It stops a memory-corruption bug, or an attacker, from overwriting a slot to redirect a call, but it only works if every slot is filled in first, so it needs eager binding. Watching a slot in gdb on a default build shows it filled before main. To see lazy binding at all, you have to ask for it with -Wl,-z,lazy.
Here's the PLT from the lazy build, objdump -d on aarch64:
00000000000005d0 <.plt>: # PLT0: the path into the loader
5d0: stp x16, x30, [sp, #-16]!
5d4: adrp x16, 1f000
5d8: ldr x17, [x16, #4088] # GOT[2]: _dl_runtime_resolve
5dc: add x16, x16, #0xff8
5e0: br x17
0000000000000630 <puts@plt>:
630: adrp x16, 20000 # page of .got.plt
634: ldr x17, [x16, #32] # load the puts slot
638: add x16, x16, #0x20 # x16 = &slot, for the resolver
63c: br x17 # jump wherever it points
0000000000000640 <main>:
...
650: bl 630 <puts@plt>Each stub is four instructions. main calls the stub at 0x630, the stub loads whatever the puts slot holds, and jumps to it. The code at 0x5d0, labelled PLT0, is the way into the loader. Step through the first call and the second:
main calls the stub, not puts itself: the compiler emitted bl 630 <puts@plt>. The puts slot in the GOT still holds the address of PLT0, the way into the loader.5.2Watching one slot get patched
Stop on the bl in gdb (the standard debugger) and step over it, and you can watch the slot change:
puts slot before the call:
0xaaaaaaac0020 <puts@got.plt>: 0x0000aaaaaaaa05d0 # PLT0, rebased
puts slot after the call:
0xaaaaaaac0020 <puts@got.plt>: 0x0000fffff7e619d0 # __GI__IO_puts in libcHere's the function that did it:
_dl_fixup (struct link_map *l, ElfW(Word) reloc_arg)
{
const ElfW(Sym) *const symtab = (const void *) D_PTR (l, l_info[DT_SYMTAB]);
const char *strtab = (const void *) D_PTR (l, l_info[DT_STRTAB]);
const uintptr_t pltgot = (uintptr_t) D_PTR (l, l_info[DT_PLTGOT]);
const PLTREL *const reloc /* which .rela.plt entry? */
= (const void *) (D_PTR (l, l_info[DT_JMPREL])
+ reloc_offset (pltgot, reloc_arg));
const ElfW(Sym) *sym = &symtab[ELFW(R_SYM) (reloc->r_info)];
void *const rel_addr = (void *)(l->l_addr + reloc->r_offset); /* the slot */
...
result = _dl_lookup_symbol_x (strtab + sym->st_name, l, &sym, l->l_scope,
version, ELF_RTYPE_CLASS_PLT, flags, NULL);
...
if (sym != NULL
&& __builtin_expect (ELFW(ST_TYPE) (sym->st_info) == STT_GNU_IFUNC, 0))
value = elf_ifunc_invoke (DL_FIXUP_VALUE_ADDR (value)); /* see 5.3 */
...
/* Finally, fix up the plt itself. */
return elf_machine_fixup_plt (l, result, refsym, sym, reloc, rel_addr, value);
}LD_DEBUG=bindings shows the order in which the bindings happen. On the lazy build, __libc_start_main binds after libc's constructors run, and puts binds after the loader prints transferring control, which means inside main:
calling init: /lib/aarch64-linux-gnu/libc.so.6
binding file ./hello-lazy [0] to libc.so.6 [0]: normal symbol `__libc_start_main' [GLIBC_2.34]
initialize program: ./hello-lazy
transferring control: ./hello-lazy
binding file ./hello-lazy [0] to libc.so.6 [0]: normal symbol `puts' [GLIBC_2.17]The full output, trimmed above, holds a surprise: hello-lazy appears to bind malloc, calloc, realloc and free, though it never calls them. That's ld.so itself. It runs on a tiny internal allocator until libc is relocated, then looks up the real malloc through the main program's scope. That's also how LD_PRELOAD gets to replace the loader's allocator.
5.3IFUNC: code that runs during relocation
memcpy in glibc isn't one function. readelf --dyn-syms on this libc marks it IFUNC, along with memmove, memset, memchr, strlen and gettimeofday. The fastest way to copy memory depends on the CPU, and libc is built before anyone knows which CPU it'll run on.
An IFUNC symbol points at a resolver. The loader calls it and uses the return value as the symbol's address. A resolver typically checks CPU features (that's what AT_HWCAP is for) and picks the SVE, SIMD or generic version of the function (SVE and SIMD are instruction sets that process several values at once).
When exactly does the resolver run? During relocation, inside the loader, while it's still patching tables, which means any library that ships an IFUNC gets to run its own code at that point. A test program makes it visible. It defines its own IFUNC sum, whose resolver records that it ran, and a constructor that checks. (The two versions of sum and the #include lines are left out.)
static int resolver_ran_before_main = 0;
static unsigned long hwcap_seen;
// Runs inside ld.so, while it processes this executable's relocations.
// On aarch64 the loader hands the resolver AT_HWCAP as its first argument.
static int (*resolve_sum(unsigned long hwcap))(const int *, int) {
resolver_ran_before_main = 1;
hwcap_seen = hwcap;
return (hwcap & HWCAP_ASIMD) ? sum_simd : sum_generic;
}
int sum(const int *a, int n) __attribute__((ifunc("resolve_sum")));
__attribute__((constructor)) static void ctor(void) {
printf("constructor: resolver already ran? %d\n", resolver_ran_before_main);
}Running it:
$ ./ifunc
constructor: resolver already ran? 1
main: sum=10, hwcap=0x40000000efb3ffff, ASIMD=yesOur resolver ran before any constructor. It was triggered by a fourth kind of relocation, R_AARCH64_IRELATIVE, which means "call this resolver and store whatever it returns". Bit 62 in that hwcap is glibc's flag saying a second argument with more feature words follows.
Resolvers run this early by design, before RELRO makes the GOT read-only. That's convenient for the loader. Section 6.4 is about what happened when someone noticed it was convenient for an attacker too.
5.4What main can count on
Now we can say what the kernel and ld.so between them promise by the time control reaches main, and what they leave open:
| Situation | Promised? | What it means |
|---|---|---|
| Every undefined symbol resolves | Yes, with eager binding. With lazy binding, a missing function is only found at its first call | With -z now the process dies before main if a symbol is missing. With lazy binding it can die halfway through a request. |
| Constructors have run | Yes | Libraries' before yours, dependencies before their dependents |
The stack holds argc, argv, envp and the auxiliary vector | Yes | The kernel's notes to user space, from section 4.2 |
| The GOT is read-only | Only with full RELRO | Needs -z now; a lazily bound program keeps a writable GOT |
Which definition of malloc you get | No | The first object in search order that exports the name wins, and LD_PRELOAD puts itself first |
| The glibc on the machine is the one you built against | No | Section 6.2 |
| How long any of this takes | No | Section 7 |
The rows that aren't a plain yes are where programs get surprised, and the next section goes through those surprises one at a time.
06When the loader surprises you
Each of these cases has a cause you can read off the binary or its environment, and most of them play out before main, so your program never gets the chance to log what went wrong.
6.1'No such file or directory', for a file that's right there
Take a copy of hello and patch the path in its .interp section (the bytes that PT_INTERP points at) to name musl's loader, the path an Alpine binary would ask for. (musl is a small alternative C library, and Alpine Linux, popular in containers, uses it. Its loader lives at /lib/ld-musl-aarch64.so.1.) Then run it on Ubuntu.
The file exists and is executable. Its PT_INTERP names /lib/ld-musl-aarch64.so.1, which isn't installed. What does execve return?
$ ls -l hello-badinterp
-rwxr-xr-x 1 root root 70312 Sep 25 13:26 hello-badinterp
$ ./hello-badinterp
bash: line 2: ./hello-badinterp: cannot execute: required file not found
$ strace ./hello-badinterp
execve("./hello-badinterp", ...) = -1 ENOENT (No such file or directory)That's the open_exec failure from section 4.1. Bash 5.2 at least says "required file", and older shells just print "No such file or directory". That plain version is what Julia Evans saw when she ran into it with a glibc binary in an Alpine container.
That case was a loader that doesn't exist. A subtler case is a loader that does exist but is too old.
6.2version `GLIBC_2.34' not found
glibc gives its symbols version tags, and a program records the versions it was linked against. Look at what hello requires:
$ objdump -T hello | grep GLIBC
0000000000000000 DF *UND* 0000000000000000 (GLIBC_2.34) __libc_start_main
0000000000000000 w DF *UND* 0000000000000000 (GLIBC_2.17) __cxa_finalize
0000000000000000 DF *UND* 0000000000000000 (GLIBC_2.17) abort
0000000000000000 DF *UND* 0000000000000000 (GLIBC_2.17) putshello calls one function, puts, which has been in glibc forever, so the GLIBC_2.17 tag is no trouble. The GLIBC_2.34 tag comes from __libc_start_main, the function your _start calls, code you didn't write. So hello, built against glibc 2.34 or newer, refuses to start on anything older. In 2.34, __libc_start_main stopped using the init argument, and the glibc source says so:
Starting with glibc 2.34, the init parameter is always NULL. Older libcs are not prepared to handle that. The macro DEFINE_LIBC_START_MAIN_VERSION creates GLIBC_2.34 alias, so that newly linked binaries reflect that dependency. (csu/libc-start.c)
So anything built on Ubuntu 22.04 or newer won't start on a glibc 2.31 or 2.28 system, and the loader reports the error before main runs. Two working rules:
- Build on the oldest system you ship to. glibc versioning is backward-compatible only: old binaries run on new glibc, never the reverse. Python's manylinux wheels are this rule turned into a standard (PEP 600).
- Or don't depend on the host's libc at all: static linking, or a container that brings its own.
Both of those failures were accidents. The next case is the loader doing exactly what it was designed to do.
6.3Someone else's malloc: interposition and LD_PRELOAD
When the loader looks up a name, the first object in the search order that defines it wins. LD_PRELOAD puts an object at the front of that order, so its malloc beats libc's for every lookup in the process, libc's own included. This is symbol interposition, and it's why a profiler or a custom allocator can slip into a program that was never built for it.
?Why does a call inside a library go through the table too?
In a normal ELF shared library, a call from f1 to f2 inside the same library still goes through a table lookup. Some other object might define f2 first, and the rules say it wins. So the call needs a relocation like any other, even though both functions are in the same file.
Let's use the hook ourselves. A working malloc counter is 30 lines of C++. It finds the real malloc with dlsym(RTLD_NEXT, "malloc"), a call that looks up the next definition of a name after the current object in the search order, and it has to cope with dlsym itself calling malloc:
// mallocount.cpp: count every malloc the program makes, print the total at exit.
#include <dlfcn.h>
#include <atomic>
#include <cstddef>
#include <cstdio>
#include <unistd.h>
using malloc_fn = void *(*)(size_t);
static malloc_fn real_malloc = nullptr;
static std::atomic<unsigned long> calls{0};
// dlsym can itself allocate. Serve those early requests from a static arena.
alignas(16) static char arena[4096];
static size_t arena_used = 0;
extern "C" void *malloc(size_t n) {
if (!real_malloc) {
static bool resolving = false;
if (resolving) { // re-entered from inside dlsym
void *p = arena + arena_used;
arena_used += (n + 15) & ~size_t(15);
return p;
}
resolving = true;
real_malloc = reinterpret_cast<malloc_fn>(dlsym(RTLD_NEXT, "malloc"));
resolving = false;
}
calls.fetch_add(1, std::memory_order_relaxed);
return real_malloc(n);
}
__attribute__((destructor)) static void report() {
char buf[64];
int len = std::snprintf(buf, sizeof buf, "[mallocount] %lu calls\n",
calls.load(std::memory_order_relaxed));
if (write(2, buf, len) < 0) {} // not stdio: it may already be torn down
}We build it as a shared library and preload it into four programs:
$ g++ -std=c++20 -O2 -shared -fPIC -o mallocount.so mallocount.cpp
$ LD_PRELOAD=$PWD/mallocount.so ./allocs # pushes 1,000 std::strings
[mallocount] 1013 calls
$ LD_PRELOAD=$PWD/mallocount.so python3 -c 'print(1)'
[mallocount] 3167 calls
$ LD_PRELOAD=$PWD/mallocount.so ./hello-static
hello
$ LD_PRELOAD=$PWD/mallocount.so ls /
$Two of the runs counted allocations, and a thousand std::strings turn out to cost 1,013 calls. The last two are the interesting ones:
- A static binary ignores
LD_PRELOADcompletely, since there's no loader to read it. lsprinted nothing.straceexplains it: coreutils closes file descriptor 2, standard error, in its exit handler, and our destructor runs after that, so itswritehas nowhere to go.
close(2) = 0
write(2, "[mallocount] 35 calls\n", 22) = -1 EBADF (Bad file descriptor)On macOS the same shim counted zero calls, with or without DYLD_FORCE_FLAT_NAMESPACE. macOS binds each import to a named library (the two-level namespace), so defining malloc somewhere else doesn't change anything. You have to say what you're replacing in a __DATA,__interpose section. With that version, the same C++ program counted 1,038 calls.
So the loader lets a library run code before main (IFUNC resolvers, constructors) and lets it change which function a name means (interposition). In 2024 someone used both to reach into ssh servers.
6.4The xz backdoor: an IFUNC resolver doing too much
On 29 March 2024 Andres Freund posted to oss-security that xz-utils 5.6.0 and 5.6.1 contained a backdoor, now CVE-2024-3094. He'd noticed ssh logins on Debian sid taking a lot of CPU, plus valgrind errors, and measured a login going from 0.299 s to 0.807 s. Everything after that is loader mechanics. Some Linux distributions patch sshd, the server that handles ssh logins, so that it calls libsystemd (systemd's library for services to report that they're ready), and libsystemd depends on liblzma, the library xz-utils builds. That chain is what the attack used:
sshd links libsystemd. OpenSSH itself doesn't use liblzma. In the GOT sits the slot for RSA_public_decrypt, which sshd calls while checking a login's signature.It only activated under narrow conditions, per Freund: x86-64 Linux, built by gcc and the GNU linker as a Debian or RPM package, argv[0] equal to /usr/sbin/sshd, TERM unset, LANG set, and LD_DEBUG and LD_PROFILE unset.
?Why didn't full RELRO stop it?
RELRO makes the GOT read-only after relocation. IFUNC resolvers run during relocation. So the attacker's code ran at a point when the tables it went after were still writable, by design.
Startup time was the clue that caught the backdoor, which is a good reason to know what normal startup costs.
07What startup costs
Everything above happens before main, so every process start pays for it. The timings in this section come from a small aarch64 Linux container. Each is a median over thousands of runs, where one run starts the program and waits for it to exit, and the rows were run interleaved so that background noise hits them all equally. The absolute numbers will differ on your machine; the gaps between rows are the part to carry away.
7.1Static, dynamic, C++, Python
| Program | Median | p10 | What's in it |
|---|---|---|---|
| hello, -static | 125 µs | 103 µs | 15 syscalls, no loader |
| hello, dynamic (-z now) | 154 µs | 133 µs | 32 syscalls, 98 symbol + 1,116 relative relocations |
| hello, dynamic, -z lazy | 152 µs | 132 µs | Same, 5 slots deferred. No difference |
| hello with std::cout | 398 µs | 346 µs | libstdc++, libm, libgcc_s. 1,794 symbol + 2,183 relative relocations |
| python3 -S -c pass | 3.4 ms | 3.2 ms | timed separately with hyperfine, 5,000 runs |
| python3 -c pass | 4.5 ms | 4.0 ms | timed separately with hyperfine, 5,000 runs |
(The "p10" column is the tenth percentile: the time that 10% of runs beat, which is close to the best case without being a lucky outlier.)
Dynamic linking costs hello about 30 µs, maybe a little less. Swapping puts for std::cout costs another 240, most of it likely loading and relocating libstdc++. Python's 3–4 ms is mostly Python itself, the interpreter's own init and imports, and -S (skip site) takes a millisecond off.
Where the file lives matters too. With both binaries in a folder shared into the container from the host machine (a bind mount, which is slow to read from), a static hello looked slower than the dynamic one: 320 µs against 186 µs. The bigger static file paid more to be paged in. Copied to local disk, the p10s were 103 µs static and 132 µs dynamic.
?Does 30 µs matter?
For a long-lived server, probably not. But a build that runs the compiler 20,000 times pays it several times per file, since gcc execs cc1, as and ld. A shell script calling grep in a loop pays it every iteration.
Hello's 98 symbol relocations are cheap. What happens when a library has tens of thousands of them, and does lazy binding save us?
7.2Lazy versus now: a crossover
To make relocation cost visible, take a generated shared library holding a chain of about 20,000 functions: f0 calls f1, which calls f2, and so on down to f20000. With default visibility every one of those internal calls goes through the PLT, since any of them could be interposed. readelf -r counts 20,002 JUMP_SLOTs. With -fvisibility=hidden and only f0 exported: 2.
main calls f0, so all 20,000 functions run. Which build starts and finishes fastest?
| Library build, what main calls | Binding | Median | Symbol relocations |
|---|---|---|---|
| default visibility, calls f19999 (2 functions run) | lazy | 216 µs | 103 |
| default visibility, calls f19999 | LD_BIND_NOW=1 | 763 µs | 20,105 |
| default visibility, calls f0 (all 20,000 run) | lazy | 1,469 µs | 20,102, by the end |
| default visibility, calls f0 | LD_BIND_NOW=1 | 1,184 µs | 20,105 |
| -fvisibility=hidden, calls f0 | either | 488 µs | 104 |
| Eager binding, 20,000 extra symbols | 763 − 216 µs | 547 µs |
| Per symbol relocation | 547 µs ÷ 20,000 | ≈ 27 ns |
| Lazy, all 20,000 called, over eager | 1,469 − 1,184 µs | 285 µs |
| Extra per lazily bound call | 285 µs ÷ 20,000 | ≈ 14 ns |
| Hidden visibility vs eager, same work | −696 µs, 2.4× faster | |
So which wins? Lazy binding wins by 3.5× when you call a handful of what you link, and loses by 24% when you call everything. Most real programs call a small fraction of libc and a large fraction of their own code, which argues for lazy. Full RELRO needs eager binding, though, and that's why distributions pay anyway.
?Why does a lazy first call cost 14 ns more than eager binding?
These numbers don't settle it. _dl_runtime_resolve stores and reloads 208 bytes of registers on every first call, which should cost a few nanoseconds, and the lookup is the same one eager binding does. A likely candidate is page faults: eager binding writes the GOT in one sequential pass, while lazy binding writes one slot at a time, scattered through the run. Running perf stat -e page-faults,dTLB-load-misses on both variants would separate the causes, but it needs hardware performance counters, which containers often hide.
Large programs with huge numbers of relocations were a real problem long before anyone timed them this way, and people have built several fixes for it.
7.3The history of trying to make this cheaper
Big C++ programs, office suites and browsers especially, are where this hurt, and the software on your machine probably carries the result:
| Fix | What it did | Where it went |
|---|---|---|
| prelink (Jakub Jelinek, Red Hat, paper) | Computed relocations ahead of time for a fixed library layout | Fedora and RHEL ran it by default. Jelinek reported "an order of magnitude difference" in relocation time for OpenOffice.org Writer (LWN, 2009). It's at odds with ASLR, since the whole point is that addresses stay put. |
| Firefox's elfhack (Mike Hommey, 2010) | Rewrote libxul's relative relocations into a compact form to cut startup I/O (glandium.org) | Retired in 2023 in favour of RELR (glandium.org) |
RELR (DT_RELR) | Standardised the compact relative-relocation format | Chrome OS carried patches from 2018, Android and Fuchsia adopted it, and it landed in glibc 2.36 (MaskRay). Linkers emit it with -z pack-relative-relocs. |
The reason so many teams bothered is size as well as time. Each ordinary relative relocation takes 24 bytes in the file, and MaskRay found relative relocations made up 7.9% of total file size across Arch Linux's /usr/bin.
All of these keep the loader and make it cheaper. Another answer is to have no loader at all.
08Static by default: Go and Alpine
Every cost and failure so far traces back to one decision, resolving addresses at run time. Some ecosystems decided otherwise, or made a different choice of C library, and you'll meet both in containers.
8.1Go: static by default
Go sidesteps GLIBC_x.y errors and missing libraries by linking statically. With cgo off (cgo is Go's bridge to C libraries), a Go binary has no PT_INTERP, so it runs even in a scratch container (an empty image with no files but yours) that has no libc at all.
One catch is that a static Go binary can't use glibc's DNS resolver, the code that turns host names into network addresses (a different "resolver" from the IFUNC kind). The net package prefers its own pure-Go resolver on Unix, and falls back to glibc's through cgo when the system's name-lookup configuration (/etc/nsswitch.conf, /etc/resolv.conf) asks for features the Go one doesn't implement. GODEBUG=netdns=go or =cgo forces one; netdns=1 logs the choice. If a Go service resolves names differently from curl on the same box, check which resolver it's using.
8.2Alpine and musl
A glibc binary on Alpine fails with the ENOENT from section 6.1, because Alpine's loader is musl's. And a musl binary behaves differently from glibc in ways musl documents: no lazy binding at all, dlclose as a no-op, and a resolver that didn't support DNS over TCP until 1.2.4 in May 2023. Before that, a truncated UDP answer just failed, so a name that resolved fine on your laptop could fail under a big DNS answer.
With the choices laid out, here's how to see what a given binary is doing and how to decide.
09Seeing inside the loader
9.1The commands
Each question this chapter raised has a command that answers it.
# What will run first, and what does this need? (sections 3 and 4)
readelf -lW ./app | grep -A1 INTERP
readelf -dW ./app | grep -E 'NEEDED|RUNPATH|RPATH|FLAGS'
objdump -T ./app | grep -o 'GLIBC_[0-9.]*' | sort -V | tail -1 # minimum glibc (section 6.2)
# RELRO status: GNU_RELRO present + BIND_NOW = full RELRO (section 5.1)
readelf -lW ./app | grep GNU_RELRO
readelf -dW ./app | grep -E 'BIND_NOW|FLAGS_1'
# The loader's own diagnostics (sections 1.2, 4.3 and 5.2)
LD_DEBUG=libs ./app # search paths, which file each soname became
LD_DEBUG=bindings ./app # every symbol, who asked, who answered
LD_DEBUG=statistics ./app # relocation counts
LD_DEBUG=help ./app # the list
LD_SHOW_AUXV=1 ./app # what the kernel handed over (section 4.2)
# Startup syscalls (section 4.3)
strace -f -e trace=openat,mmap,mprotect ./app9.2Rules that hold up
- Build shared libraries with
-fvisibility=hiddenand export an explicit API. Smaller file, fewer lookups, faster calls. It was the biggest single win in section 7.2. - Keep full RELRO (
-z relro -z now). Ubuntu's default already does. It costs eager binding, and in exchange an attacker can't overwrite GOT entries after startup. - Link RELR with
-z pack-relative-relocsif your toolchain and target glibc (2.36 or newer) support it. - Count your libraries. Each
NEEDEDis anopenat, severalmmaps, relocations and its constructors.lddon a large C++ service can easily list dozens. - Go static for tools you exec a lot, if you can live with the trade-offs in the next table.
9.3What you trade for what
| Choice | You get | You pay |
|---|---|---|
| Dynamic linking | Shared pages, security fixes without rebuilds, LD_PRELOAD | ~30 µs per exec for hello, far more for C++; GLIBC_x.y errors |
| Static linking | One file, no loader, no version skew | No LD_PRELOAD, rebuild for every libc CVE, and glibc's NSS/dlopen features break |
| Lazy binding | 3.5× faster startup when few symbols are called | Writable GOT for the process lifetime; a missing symbol is a crash mid-run |
Full RELRO (-z now) | Read-only GOT after startup | Every symbol bound up front, used or not |
-fvisibility=hidden | Fewer relocations, direct calls, smaller file | You must mark the API explicitly; nobody can interpose internals |
| musl (Alpine) | Small, simple, static-friendly | No lazy binding, dlclose is a no-op, DNS over TCP only since 1.2.4 |
9.4Symptom, cause, fix
| Symptom | Likely cause | Fix |
|---|---|---|
| "No such file or directory" (or "required file not found") for a file that exists | PT_INTERP names a loader that isn't installed, often a glibc binary on Alpine | readelf -l ./bin | grep interpreter; build for the target's libc |
version `GLIBC_2.34' not found | Built against a newer glibc than the target has | Build on the oldest target, use a manylinux-style image, or link statically |
cannot open shared object file | The loader can't find a NEEDED soname | readelf -d for NEEDED and RUNPATH; LD_DEBUG=libs for the search |
LD_PRELOAD has no effect | The binary is static, or it's macOS's two-level namespace | Check for PT_INTERP; on macOS use __DATA,__interpose |
| A preload shim crashes or hangs at startup | dlsym allocating and re-entering your malloc | Serve re-entrant calls from a static arena |
| Slow start for a big C++ service | Tens of thousands of symbol relocations | -fvisibility=hidden, fewer libraries, RELR |
| A Go service resolves names differently from curl | The pure-Go resolver | GODEBUG=netdns=1 to see which, =go or =cgo to force one |
10Summary
- A compiled program calls functions whose code isn't in its file, so someone has to supply their addresses at run time. Static linking copies the code in; dynamic linking shares one copy and fills in addresses at start-up.
- The CPU starts in ld.so, not your program. The kernel maps both, reads
PT_INTERP, and hands your entry point over asAT_ENTRY. - An ELF file describes itself twice. Program headers are for loading; section headers are for tools, and the kernel never reads them.
- A relocation is "write this address here". Relative ones are cheap; symbol ones need a lookup through every loaded object.
- GNU hash makes "not here" cheap. A Bloom filter rejects most objects with one load and two shifts.
- Ubuntu binds everything at startup by default.
-z nowgives full RELRO and a read-only GOT; lazy binding needs-z lazy. - IFUNC resolvers run inside the loader, before RELRO. That's the door the xz backdoor used to reach
sshd. - "No such file or directory" usually means the interpreter. The kernel returns
ENOENTfor the loader that isn't there. - Even hello needs GLIBC_2.34 if you build it on a recent distribution, so build on the oldest system you ship to.
LD_PRELOADwins every lookup, and static binaries ignore it.- Hidden visibility beats both binding modes. 488 µs against 1,184 for the same 20,000 calls, and a library half the size.
11Build this
Write an LD_AUDIT library that times every binding in a real program.
glibc's rtld-audit(7) interface calls your la_objopen for every library loaded and la_symbind64 for every symbol bound. That's more or less the same machinery the xz payload hooked.
- Implement
la_version,la_objopenandla_symbind64, and log object name, symbol name and aclock_gettimetimestamp for each. - Run it on
python3 -c passand on a C++ service you own. Count bindings per library. - Then rebuild one of your own libraries with
-fvisibility=hiddenand watch its bindings disappear from the log.
Your first run will probably show you a library you didn't know was loaded. On Debian and Ubuntu, liblzma in sshd was exactly that.
12Interview questions
beginnerWhat's the difference between program headers and section headers?›
Program headers describe segments: what to map, where, with which permissions.
That's all the kernel and the dynamic loader read. Section headers describe
the file for tools like the linker, objdump and gdb: .text, .dynsym,
.rela.plt. Linux never looks at them, so a binary can run with its
section header table stripped.
beginnerA binary exists and is executable, and running it says 'No such file or directory'. Why?›
Almost always the interpreter named in PT_INTERP doesn't exist, typically a
glibc binary in an Alpine container asking for /lib/ld-linux-*.so. Linux
opens that path inside load_elf_binary and returns its ENOENT from
execve. readelf -l ./bin | grep interpreter shows what it wants.
intermediateWalk through the first call to puts in a lazily bound program.›
main calls puts@plt. That stub loads the puts slot from .got.plt. It
still points at PLT0, and jumps there. PLT0 pushes the slot address and jumps to
_dl_runtime_resolve. That saves the argument registers and calls _dl_fixup.
That finds the .rela.plt entry, looks puts up through the GNU hash tables of
each loaded object, writes the address into the slot, and returns it. The
trampoline restores registers and jumps to puts. Later calls skip all of it.
intermediateWhy does a program built on Ubuntu 24.04 fail with GLIBC_2.34 not found on an older box?›
glibc versions its symbols, and a program records the version it linked
against. Even a hello world references __libc_start_main@GLIBC_2.34, from
crt1.o, because 2.34 changed how that function handles constructors. Older
glibc doesn't have that version node, so the loader refuses before main.
Build on the oldest target, use a manylinux-style build image, or link
statically.
intermediateWhat does full RELRO protect, and what does it cost?›
It makes the GOT and other relocated data read-only once the loader finishes, so
a memory-corruption bug can't redirect a function pointer in the GOT. It
requires binding every symbol at startup (-z now). In the 20,000-symbol
test that cost 547 µs against lazy binding when only two functions were called.
deepHow did the xz backdoor get code running inside sshd, and why didn't RELRO help?›
Distro-patched sshd links libsystemd, which links liblzma. That malicious
liblzma replaced its IFUNC resolvers. The loader calls those during relocation,
before main and before RELRO is applied. From there it installed a
dynamic-linker audit hook and redirected RSA_public_decrypt when that symbol
was bound. RELRO locks tables after relocation, and the attack happened during it.
deepWhy does -fvisibility=hidden make a shared library start faster?›
Default visibility means every exported function can be interposed by an
earlier object, so even internal calls go through the PLT and need a
JUMP_SLOT relocation. Hidden symbols can't be interposed, so the static
linker emits direct branches and drops them from .dynsym. In the test that took
a library from 20,002 jump slots to 2, from 3.7 MB to 2.0 MB, and cut startup
from 1,184 µs to 488.
deepYour LD_PRELOAD malloc shim deadlocks or crashes at startup. What's the likely cause?›
Probably recursion during bootstrap. Your shim's malloc calls dlsym(RTLD_NEXT, "malloc")
to find the real one, and dlsym itself can allocate, and that re-enters the shim
before the pointer is set. The usual fix is a static arena for re-entrant calls
during resolution, or a glibc-specific entry point. Printing from a destructor
has its own trap: some programs close stderr before destructors run.
13Go deeper
Where does the CPU start executing after execve of a dynamic PIE?›
In ld.so, at the interpreter's entry point. The program's own entry goes on the stack as AT_ENTRY and ld.so jumps there after relocation.
Why did LD_PRELOAD do nothing to the static binary?›
There's no dynamic loader to read the variable. Static binaries have no PT_INTERP and resolve everything at link time.
Lazy binding or BIND_NOW: which starts faster?›
It depends on how much you call. Lazy was 3.5× faster when two of 20,000 functions ran, and 24% slower when all of them did.
Your hello world was built on Ubuntu 24.04 with gcc -O2. Is it lazily bound?›
No. Ubuntu's toolchain links with -z now by default, so FLAGS shows BIND_NOW and the GOT is read-only before main.
How fork() and exec() start a new program, the step just before
everything in this chapter. It doesn't cover ELF or the loader, so it pairs
with sections 4 and 5 here.
Free online.
The paper on relocation cost, visibility and GNU hash, from glibc's former maintainer. Old and still correct on every mechanism in section 3.
load_elf_binary and create_elf_tables. About 500 lines, and all of sections 4.1 and 4.2.
dl-runtime.c,
dl-lookup.c
and the aarch64
dl-trampoline.S.
Read _dl_fixup first.
The history from elfhack to DT_RELR, with size measurements across a whole distribution.
14Related chapters
Chapter 04. Every mmap in the strace above
creates a mapping whose pages fault in on first touch. That's part
of why a binary on a bind mount started so much slower than the same binary
on local disk.
Chapter 07. execve is the
heaviest syscall most programs make, and AT_SYSINFO_EHDR in the auxv is
where the vDSO from that chapter gets handed over.
Chapter 05. What you're replacing when an
LD_PRELOAD of jemalloc or tcmalloc interposes malloc.
Chapter 06. The fork half of
fork-and-exec, and what the new process inherits before the loader runs.