KnowSys

Linking, Loading & How a Process Starts

Follow one tiny program from the moment you press Enter to the first line of main: how the kernel opens the file, how a second program called the dynamic loader finds the libraries your program calls, and what all of that costs every time anything starts.

⏱ 42 min read◆ IntermediateAssumes: a terminal, and Docker if you want to run the commands; chapters 04 (virtual memory) and 07 (syscalls) help
Start reading

You write a four-line C program called hello.c that calls puts("hello"), compile it with gcc -O2 -o hello hello.c, type ./hello and press Enter. "hello" appears and the program exits, and the whole trip from Enter to exit takes about 150 microseconds. It's natural to picture the compiled file as the entire program: a block of machine code that the computer copies into memory and starts running.

That picture leaves out puts. Your file contains a call to it, but the code that writes to the screen lives in the C library, a separate file on disk that gcc never copied in. When you press Enter, something has to find that file, bring it into memory and tell your program where puts ended up, all before your first line runs. If any of that fails, the program dies before it begins, and the error message often points at the wrong thing.

A program called the dynamic loader, ld.so, does this work, and the kernel starts it before your own code gets a turn. This chapter follows one run of hello from Enter to the first line of main, and for each step asks what it does, why it has to exist, what it costs and how it fails. We start by asking a real program what it needs, work out why that need exists, read the file that states it, and then watch the whole sequence happen.

01What a program needs before it runs

1.1Asking ls what it needs

Let's begin with something real. ls is small and has run on your machine thousands of times, so it makes a fair test subject. These commands run inside a throwaway Ubuntu 24.04 Docker container so that everyone sees the same files, but any Linux machine gives the same kind of answer.

The container runs three commands, and each needs a word of explanation.

The first, ldd /bin/ls, lists the shared libraries that ls needs. A shared library is a file of ready-made functions that many programs use. The most important one is libc.so.6, the C library, which holds puts, malloc, read and thousands of other functions.

The second runs ls with LD_DEBUG=libs in front of it. That's an environment variable, a named setting that a program inherits from the shell that started it. Writing NAME=value command sets it for that one command, and this particular one asks the loader to describe its search for each library. The loader writes that description to standard error, so 2>&1 >/dev/null sends standard error down the pipe and throws away ls's normal output, and head -12 keeps the first dozen lines.

The third sets LD_PRELOAD, another environment variable, which names a library to load before all the others. We give it a file that doesn't exist, to see what the loader does about it.

List the libraries /bin/ls needs, watch the loader search for them, and give it a bad one
shell
Shell
docker run --rm ubuntu:24.04 sh -c '
  ldd /bin/ls
  echo
  LD_DEBUG=libs /bin/ls / 2>&1 >/dev/null | head -12
  echo
  LD_PRELOAD=/nope.so /bin/ls / 2>&1 | head -3'
output
C++
	linux-vdso.so.1 (0x0000ffff9aa17000)
	libselinux.so.1 => /lib/aarch64-linux-gnu/libselinux.so.1 (0x0000ffff9a940000)
	libc.so.6 => /lib/aarch64-linux-gnu/libc.so.6 (0x0000ffff9a780000)
	/lib/ld-linux-aarch64.so.1 (0x0000ffff9a9da000)
	libpcre2-8.so.0 => /lib/aarch64-linux-gnu/libpcre2-8.so.0 (0x0000ffff9a6d0000)
 
        14:	find library=libselinux.so.1 [0]; searching
        14:	 search cache=/etc/ld.so.cache
        14:	  trying file=/lib/aarch64-linux-gnu/libselinux.so.1
        14:	
        14:	find library=libc.so.6 [0]; searching
        14:	 search cache=/etc/ld.so.cache
        14:	  trying file=/lib/aarch64-linux-gnu/libc.so.6
        14:	
        14:	find library=libpcre2-8.so.0 [0]; searching
        14:	 search cache=/etc/ld.so.cache
        14:	  trying file=/lib/aarch64-linux-gnu/libpcre2-8.so.0
        14:	
 
ERROR: ld.so: object '/nope.so' from LD_PRELOAD cannot be preloaded (cannot open shared object file): ignored.
bin
boot

A tiny program like ls depends on three libraries: libselinux, libc and libpcre2. The list also holds the loader itself, /lib/ld-linux-aarch64.so.1, and linux-vdso.so.1, which isn't a file on disk at all. It's a small piece of code the kernel places in every process's memory so that a few system calls can run without entering the kernel (chapter 07 calls it the vDSO). Each line ends with an address, the place in memory where that piece landed for this run. Those numbers differ on every run, and section 2 explains why.

1.2What the second and third command show

In the second block the loader narrates its search. For each library it first checks a cache file, /etc/ld.so.cache, a prebuilt index from library names to file paths that saves it from searching directories one by one. Then it tries the path the cache gave it. The 14 at the start of each line is the process ID, the number the kernel gave this run of ls.

The third block shows LD_PRELOAD at work. It's a hook into the loader: any library you name there gets loaded before all the others, and its functions win over the real ones. A file that doesn't exist only produces a warning, and ls goes on to list / without it. That hook is how profilers and replacement allocators such as jemalloc get into a program without rebuilding it, and section 6 shows what it can do.

Everything that follows explains what happened between you pressing Enter and ls printing its first line: how the kernel starts the loader, what's in the file it reads, and what it costs every time any process starts.

1.3The program we'll follow

ls already needs three libraries, which is more than we want to draw. For the rest of the chapter we'll follow hello, the smallest program that still has the problem. Here is the source, and what ldd says about the compiled file:

C
// hello.c
#include <stdio.h>
int main(void) { puts("hello"); return 0; }
Output
$ gcc -O2 -o hello hello.c
$ ./hello
hello
$ ldd ./hello
    linux-vdso.so.1 (0x0000ffff96b09000)
    libc.so.6 => /lib/aarch64-linux-gnu/libc.so.6 (0x0000ffff968e0000)
    /lib/ld-linux-aarch64.so.1 (0x0000ffff96acc000)

hello needs only libc, plus the loader and the vDSO page that every process gets. The #include <stdio.h> line only declares puts, telling the compiler what its arguments look like. It doesn't contain its code. So the compiled hello holds a call to a function whose code is in another file, and the obvious question is where the address for that call comes from.

02Why the program doesn't contain puts

2.1A call to an address nobody knows

A recipe card says "add the tomato sauce from the pantry" and "use the stock on shelf two". It doesn't contain the sauce or the stock. Before anyone cooks, an assistant reads the card, walks to the pantry, fetches every item it names and lays them on the counter. Your program is the recipe card, and the loader is the assistant.

Here is how the card gets written. gcc turns your call to puts into a branch instruction, bl puts on the ARM processors these examples run on, which means "jump to the address of puts and come back afterwards". puts is a symbol, a name for a function or variable that one piece of code defines and another uses. But the compiler only ever saw hello.c. It never saw the C library's code, so it can't write down the address of puts. Somebody has to supply that number, and there are two ways to do it: at build time, or at run time.

2.2Static linking: copy the code in

Doing it at build time is the simplest answer. A tool called the linker (ld, which gcc runs for you at the end of a build) combines the compiled pieces into one executable. It can copy the machine code for puts, and everything puts calls, into your executable and patch the branch to point at the copy. This is static linking, and it works. If you build hello with gcc -static, the result needs no libraries: ldd reports "not a dynamic executable".

Flowchart: a library and two object files feed into the linker, which can produce a library, a DLL or an executable
The linker's job in one picture: compiled object files and libraries go in, and an executable or a library comes out. The labels are Windows names. On Linux the shared library is a .so file, and the executable is an ELF file like hello.Image: Qef, public domain, via Wikimedia Commons

It also makes the file much bigger. Static hello is 639,808 bytes, against 70,312 for the dynamic one.

?Why not do that for every program?

Think about what it means across a whole system. Every program on the disk carries its own copy of the C library, and every running process holds its own copy in RAM. And when a security fix lands in the library, every program has to be relinked to pick it up.

2.3Dynamic linking: one copy, many processes

So most systems share instead. One copy of libc.so.6 sits on disk, and each process that needs it asks the kernel to make the file's bytes appear at some range of the process's memory addresses. This is called mapping the file. The kernel manages memory in pages, chunks of 4 KB on most systems, and when many processes map the same file, it keeps one copy of each page in physical memory and shows that copy to all of them. A thousand processes then cost roughly one copy of libc. Sharing one library file this way, and connecting each program to it when the program starts, is dynamic linking, and it creates a new problem.

hello can't have the address of puts written into it, because libc lands somewhere different in each process. That's partly on purpose. Linux uses ASLR (address space layout randomisation): it places code at random addresses on every run, so an attacker who has found a bug can't rely on knowing where any function lives. So nobody knows the address of puts until the process starts.

The fix is a level of indirection. hello's call goes through an entry in a table inside the program's own memory (we'll meet it as the GOT in section 3), and the table starts out empty. Before main runs, someone has to find each library by name, map it, work out where each function landed, and write those addresses into the table. That someone is the dynamic loader, ld.so, which the kernel starts before your program gets control. Each fill-in is a relocation, an instruction of the form "write this address in this spot". Starting hello takes 1,214 of them, by the loader's own count. Only a handful belong to hello itself, and almost all the rest are libc fixing up pointers to its own code for wherever it landed.

Static linkingDynamic linking
When addresses are filled inAt build time, by the linkerAt process start (or first call), by ld.so
Copies of libcOne per program, on disk and in RAMOne file, shared pages
A libc security fixRelink every programReplace one file
Who runs firstYour _startld.so

The loader can't guess any of this. Each file has to carry its own instructions, so the next question is how a file can say "I need libc, and here are the spots to fix".

03Inside an ELF file

Linux stores executables and libraries in a format called ELF, the Executable and Linkable Format. An ELF file is a header, two tables that describe the rest of the file, and the data those tables point at. We'll open hello and read what the loader reads.

3.1The header, and the 64 bytes the kernel reads first

Every ELF file opens with a fixed header, 64 bytes long on 64-bit targets. The kernel checks the magic number, the first four bytes, 7f 45 4c 46, which spell \x7fELF. Then it checks the file's type and the processor it's for. The header also holds two offsets, positions in the file where the two tables start: e_phoff for the program headers and e_shoff for the section headers. Section 3.2 explains why there are two.

Poster dissecting a tiny ARM ELF file byte by byte: ELF header, program header table, code, data, section names and section header table, with the loading process described below
Ange Albertini's walk through a complete, tiny ELF file, a 32-bit ARM program that prints Hello World. At the top right, the first bytes are 7f 45 4c 46, followed by fields including e_entry and e_phoff. Being 32-bit, its header is 52 bytes instead of 64, but the fields are the same. The loading steps along the bottom say it plainly: sections are not used.Image: Ange Albertini, CC BY 1.0, via Wikimedia Commons
0234566.58e_ident16btype/machine/version8be_entry8be_phoff8be_shoff8bsizes & counts12bflags 4b7f 45 4c 46 = \x7fELF0x680sections @68520DYN, AArch64program headers @649 phdrs, 28 sections
The ELF64 header, each field drawn in proportion to its size. The kernel uses e_entry and e_phoff; it never follows e_shoff to the section headers.

Those numbers are from a hello built with plain gcc -O2 on Ubuntu 24.04. Its type field says DYN, the type used for shared libraries, where you might expect EXEC, the type for a program that must sit at one fixed address. That's because a default Ubuntu build is a position-independent executable (PIE): its code works wherever in memory it's placed. A PIE lets the kernel put the program itself at a random address, the same ASLR it applies to libraries, and Linux treats it almost exactly like a shared library. e_entry, 0x680 here, is where execution begins, counted from wherever the file ends up being placed.

3.2Program headers are for loading, section headers are for tools

An ELF file describes itself twice, because two different readers want different things. The kernel and the loader only need to know what to put in memory, where, and with which permissions. A linker or debugger wants to know what each stretch of bytes means.

TableDescribesRead by
Program headers (segments)What to map, where, and with which permissionsThe kernel and ld.so
Section headersWhat the bytes mean: .text (the code), .dynsym (the dynamic symbols), .rela.plt (the function-call relocations)The linker, objdump, gdb
Diagram of an ELF file: ELF header, program header table, sections .text, .rodata and .data, and a section header table, with arrows from the program header table to groups of sections and from the section header table to each section
One file, described twice. Arrows from the program header table point at whole segments, each a run of sections mapped as one piece. Arrows from the section header table point at the individual sections inside them, such as .text, .rodata and .data.Image: Surueña, CC BY-SA 3.0, via Wikimedia Commons

A segment is a stretch of the file that gets mapped as one piece. Here are hello's segments:

Output
$ readelf -lW hello          # trimmed
  Type           Offset   VirtAddr           FileSiz  MemSiz   Flg Align
  PHDR           0x000040 0x0000000000000040 0x0001f8 0x0001f8 R   0x8
  INTERP         0x000238 0x0000000000000238 0x00001b 0x00001b R   0x1
      [Requesting program interpreter: /lib/ld-linux-aarch64.so.1]
  LOAD           0x000000 0x0000000000000000 0x0008b4 0x0008b4 R E 0x10000
  LOAD           0x00fd90 0x000000000001fd90 0x000280 0x000288 RW  0x10000
  DYNAMIC        0x00fda0 0x000000000001fda0 0x0001f0 0x0001f0 RW  0x8
  GNU_STACK      0x000000 0x0000000000000000 0x000000 0x000000 RW  0x10
  GNU_RELRO      0x00fd90 0x000000000001fd90 0x000270 0x000270 R   0x1

There are two LOAD segments, which are the ones the loader maps. One is marked R E (readable and executable: the code) and the other RW (readable and writable: the data). No segment is both writable and executable, so data can't be run as code. INTERP names the program that runs first, which will matter a great deal in section 4. DYNAMIC points at the loader's instructions, covered next. GNU_RELRO marks the part of the writable segment that goes read-only once relocation is done, and section 5 explains why that's worth doing.

?Why is this file 70 KB for 2 KB of code?

Because of that 0x10000 in the Align column. A segment can only be mapped starting at a page boundary, and ARM Linux kernels can be built with 4 KB, 16 KB or 64 KB pages. The aarch64 toolchain aligns segments to 64 KB so the same binary works on all three, and the file is padded to match.

3.3.dynamic, relocations, and the GOT

The ELF specification names segment types with a PT_ prefix, which readelf leaves off, so the DYNAMIC line above is PT_DYNAMIC and INTERP is PT_INTERP. PT_DYNAMIC points at the .dynamic section, an array of tag and value pairs that works as the loader's table of contents. Hello's has 27 entries. These are the ones that matter:

TagWhat it tells the loader
NEEDED libc.so.6The dependency list, by soname, a library's name with a version number in it so an incompatible new version gets a different name
SYMTAB, STRTAB, GNU_HASHThe dynamic symbol table, its strings, and the hash table for looking names up
RELA, JMPRELTwo relocation tables: .rela.dyn for data, .rela.plt for function calls
PLTGOTThe address of the GOT, the Global Offset Table: the table of addresses that hello's calls go through
FLAGS BIND_NOW, FLAGS_1 NOW PIEBind everything at startup. Section 5 comes back to this.

A relocation is an instruction to the loader: at this offset, write the address of that symbol, plus this addend. Here's readelf -r on a lazily linked build of the same hello (section 5 explains "lazily"):

Output
Relocation section '.rela.dyn' at offset 0x480 contains 8 entries:
    Offset             Type               Symbol's Name + Addend
000000000001fdc8  R_AARCH64_RELATIVE                        790
000000000001ffc8  R_AARCH64_GLOB_DAT     __cxa_finalize@GLIBC_2.17 + 0
...
Relocation section '.rela.plt' at offset 0x540 contains 5 entries:
0000000000020000  R_AARCH64_JUMP_SLOT    __libc_start_main@GLIBC_2.34 + 0
0000000000020020  R_AARCH64_JUMP_SLOT    puts@GLIBC_2.17 + 0

That's three kinds of relocation, with three costs:

TypeWhat the loader doesCost
RELATIVEAdds the load base and storesNo lookup at all
GLOB_DATLooks up a data or function symbolA hash, a walk over every loaded object, a string compare
JUMP_SLOTLooks up a function and fills its GOT slotSame lookup, but it can be put off until the first call

RELATIVE relocations are pointers within the same file, such as "this variable holds the address of that function", so the loader only has to add the address where the file landed. The other two name a symbol, and finding the symbol is where the time goes.

3.4GNU hash: a Bloom filter in front of every lookup

A symbol lookup asks each loaded object in turn, "do you define puts?" (A loaded object is the executable or any library the loader has mapped.) Most of them say no, so making "no" cheap is what matters. DT_GNU_HASH does this with a Bloom filter, a small array of bits that can say "definitely not here" after testing just a couple of bits, and sometimes says "maybe", in which case the real table has to be checked.

A row of bits, with three items x, y and z each setting three bits to 1, and an item w whose three bits include a 0
A Bloom filter in general. Adding x, y and z each set a few bits, chosen by hashing the name. To test w, hash it the same way and check its bits: one of them is 0, so w is definitely not in the set. If all its bits had been 1, the answer would only be maybe. GNU hash tests two bits per name.Image: David Eppstein, public domain, via Wikimedia Commons

Here is one lookup of puts against hello's executable and then libc:

Looking up puts, one object at a time
ld.soasks each object in search orderhelloBloom filter, bucketslibc.so.6Bloom filter, buckets, chainsputshash computedBloom filterBloom filterbuckethash % nbucketschain entry31-bit comparedo you define it?
Step 1. The loader needs the address of puts. It hashes the name once, then asks each object in search order: the executable first, then libc.
1 / 6

Here is the same lookup in glibc 2.39's source. Two details the picture leaves out: the two bits come from different parts of the hash (hashbit1 from its low bits, hashbit2 after shifting it right), and the lowest bit of each chain entry marks the end of the chain.

C
const ElfW(Addr) *bitmask = map->l_gnu_bitmask;
if (__glibc_likely (bitmask != NULL))
  {
    ElfW(Addr) bitmask_word
      = bitmask[(new_hash / __ELF_NATIVE_CLASS)
                & map->l_gnu_bitmask_idxbits];
 
    unsigned int hashbit1 = new_hash & (__ELF_NATIVE_CLASS - 1);
    unsigned int hashbit2 = ((new_hash >> map->l_gnu_shift)
                             & (__ELF_NATIVE_CLASS - 1));
    /* Two bits must both be set, or this object can't define the name. */
    if (__glibc_unlikely ((bitmask_word >> hashbit1)
                          & (bitmask_word >> hashbit2) & 1))
      {
        Elf32_Word bucket = map->l_gnu_buckets[new_hash % map->l_nbuckets];
        if (bucket != 0)
          {
            const Elf32_Word *hasharr = &map->l_gnu_chain_zero[bucket];
            do
              /* Compare 31 bits of hash before any strcmp. */
              if (((*hasharr ^ new_hash) >> 1) == 0)
                { /* ... check_match(): the actual string compare ... */ }
            while ((*hasharr++ & 1u) == 0);   /* low bit ends the chain */
          }
      }

Note the __glibc_unlikely on the Bloom test. glibc's authors expect most objects to answer "no", so the compiler is told to lay out the code with that case as the fast path. In a C++ program with a dozen libraries, most probes probably end at one 64-bit load and two shifts. For hello's two objects the filter barely matters, and it pays off as the number of loaded objects grows.

That completes what the file says: which libraries it needs, which spots to fix, and how to find names quickly. What's left is to watch who reads all this, and in what order, when you press Enter.

04From execve to main

When you press Enter, the shell asks the kernel to start hello with the system call execve (the shell has first made a copy of itself with fork, and the copy calls execve). That call replaces the copy's program with the one in the file. Everything between the kernel taking over and your first line of main is the path we want to understand. Here is the whole thing for a dynamically linked hello, with the three players in it: the kernel, the loader and your code.

One execve, up to main
Files on diskhello, ld.so, libc.so.6KernelexecveThe new processits own address spacehellothe ELF fileld.sothe loaderlibc.so.6the C libraryexecve./hellohellomappedld.somappedauxvkernel's noteslibc.so.6mappedread headers
Step 1. The shell calls execve("./hello"). The kernel opens the file and reads its program headers. One of them, INTERP, names another program that must run first: the loader.
1 / 7

Next we take the frames one at a time, starting with the kernel's part. One thing to notice already is that hello and libc were never "loaded" by copying. Mapping makes their bytes appear in the process, and the pages are brought in from disk only when something touches them.

4.1The kernel: load_elf_binary

execve walks a list of binary formats until one claims the file. For ELF that's load_elf_binary in fs/binfmt_elf.c. It reads the program headers, and the first thing it hunts for is PT_INTERP:

C
for (i = 0; i < elf_ex->e_phnum; i++, elf_ppnt++) {
    char *elf_interpreter;
    ...
    if (elf_ppnt->p_type != PT_INTERP)
        continue;
 
    retval = -ENOEXEC;
    if (elf_ppnt->p_filesz > PATH_MAX || elf_ppnt->p_filesz < 2)
        goto out_free_ph;
 
    retval = -ENOMEM;
    elf_interpreter = kmalloc(elf_ppnt->p_filesz, GFP_KERNEL);
    ...
    retval = elf_read(bprm->file, elf_interpreter, elf_ppnt->p_filesz,
                      elf_ppnt->p_offset);
    ...
    interpreter = open_exec(elf_interpreter);   /* the path from .interp */
    kfree(elf_interpreter);
    retval = PTR_ERR(interpreter);
    if (IS_ERR(interpreter))
        goto out_free_ph;                        /* ENOENT comes from here */

So the path inside .interp gets opened by the kernel, from the kernel's view of the filesystem. Keep that last comment in mind for section 6.1.

Next it maps each PT_LOAD segment at a random base address (a PIE has type ET_DYN, the DYN from section 3.1, so it gets ASLR), maps the interpreter the same way, and picks where to jump. In the code below, load_bias is that random base:

C
e_entry = elf_ex->e_entry + load_bias;          /* your program's entry */
...
if (interpreter) {
    elf_entry = load_elf_interp(interp_elf_ex, interpreter,
                                load_bias, interp_elf_phdata, &arch_state);
    if (!IS_ERR_VALUE(elf_entry)) {
        interp_load_addr = elf_entry;
        elf_entry += interp_elf_ex->e_entry;    /* ...but we start in ld.so */
    }
    ...
} else {
    elf_entry = e_entry;                        /* static: straight to _start */
}
...
retval = create_elf_tables(bprm, elf_ex, interp_load_addr,
                           e_entry, phdr_addr);
...
START_THREAD(elf_ex, regs, elf_entry, bprm->p);

So there are two entry points. The CPU starts in ld.so, and your program's own entry goes on the stack as AT_ENTRY for ld.so to jump to later.

4.2The auxiliary vector: the kernel's notes

create_elf_tables writes the auxiliary vector onto the stack, next to argv and envp. It's a list of facts the kernel already knows and the loader would otherwise have to work out, such as where the vDSO is, how big a page is, and where hello's program headers ended up. glibc will print it for you:

Output
$ LD_SHOW_AUXV=1 ./hello        # trimmed
AT_SYSINFO_EHDR:      0xffffbc9dc000      # the vDSO from chapter 07
AT_HWCAP:             efb3ffff            # CPU features (section 5.3)
AT_PAGESZ:            4096
AT_PHDR:              0xaaaac8e80040      # where your program headers ended up
AT_BASE:              0xffffbc99f000      # where ld.so was mapped
AT_ENTRY:             0xaaaac8e80680      # 0x680 from the header, plus the bias
AT_RANDOM:            0xffffe499fd78      # 16 bytes: stack canary, pointer guard

AT_HWCAP is a bitmask of CPU features, which section 5.3 uses. AT_RANDOM points at 16 random bytes that libc turns into secret values: a "canary" placed on the stack so a buffer overflow that overwrites it gets caught, and a key for scrambling saved function pointers. AT_ENTRY is the 0x680 from the ELF header plus the random base, which is hello's _start. ld.so doesn't have to parse your ELF file again. Linux already did, and this is it handing over its notes.

4.3ld.so: map the libraries, apply the relocations

Now the loader is running. Here's the complete strace of the dynamic hello, with only arguments trimmed. (strace prints every system call a program makes. openat opens a file, mmap maps a file or a range of memory into the process, and mprotect changes the permissions on a range that's already mapped.)

Output
execve("./hello", ["./hello"], 0xffffc129ca80 /* 6 vars */) = 0
brk(NULL)                               = 0xaaaadd59a000
mmap(NULL, 8192, PROT_READ|PROT_WRITE, ...)
faccessat(AT_FDCWD, "/etc/ld.so.preload", R_OK) = -1 ENOENT
openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
fstat(3, ...); mmap(NULL, 10999, PROT_READ, MAP_PRIVATE, 3, 0); close(3)
openat(AT_FDCWD, "/lib/aarch64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
read(3, "\177ELF\2\1\1\3\0\0\0\0\0\0\0\0\3\0\267\0..., 832) = 832
mmap(NULL, 1892240, PROT_NONE, ...)     # reserve the whole span first
mmap(0xffffb6ca0000, 1826704, PROT_READ|PROT_EXEC, MAP_FIXED, 3, 0)
munmap(...); munmap(...); mprotect(..., PROT_NONE)
mmap(0xffffb6e4d000, 20480, PROT_READ|PROT_WRITE, MAP_FIXED, 3, 0x19d000)
mmap(0xffffb6e52000, 49040, PROT_READ|PROT_WRITE, MAP_FIXED|MAP_ANONYMOUS)
close(3)
set_tid_address(...); set_robust_list(...); rseq(...)
mprotect(0xffffb6e4d000, 12288, PROT_READ)    # libc's RELRO
mprotect(0xaaaad2a2f000, 4096, PROT_READ)     # hello's RELRO
mprotect(0xffffb6eaa000, 8192, PROT_READ)     # ld.so's RELRO
...
write(1, "hello\n", 6)                  = 6
exit_group(0)                           = ?

Read it in order. After execve, the loader looks for /etc/ld.so.preload and opens /etc/ld.so.cache, its first lookups. It then opens libc.so.6, reads the first 832 bytes (the ELF header and program headers) and makes a run of mmap calls, one per segment. The three mprotect(..., PROT_READ) calls near the end are full RELRO being applied to libc, to hello and to ld.so itself, which section 5 explains. The write is hello's own output, and exit_group ends the process.

That's 32 syscalls in total. A static build makes 15.

?Why does nobody open hello or ld.so?

Because the kernel mapped both already, in load_elf_binary. The loader opens only what the kernel didn't: the cache and the libraries.

Two more things stand out:

  • The loader checks /etc/ld.so.preload before it reads the cache. It's a system-wide LD_PRELOAD that persists across reboots, and a favourite of rootkits, malware that hides itself by replacing the functions tools like ls use.
  • It reserves libc's whole address range as PROT_NONE first, then maps the segments into it with MAP_FIXED, so nothing else can land in the gap between them.

Between the last close(3) and the first mprotect is where relocations get applied, with no syscalls to see. LD_DEBUG=statistics counts them:

Output
$ LD_DEBUG=statistics ./hello
      3992:	                 number of relocations: 98
      3992:	      number of relocations from cache: 7
      3992:	        number of relative relocations: 1116

These are the 1,214 from section 2.3. Almost all of the 1,116 relative ones are libc fixing up its own pointers, and the 98 symbol relocations are the ones that cost lookups. On some architectures, x86-64 among them, glibc also prints how many CPU cycles startup took; the aarch64 build of glibc prints no such line.

4.4Constructors, __libc_start_main, and finally main

After relocation ld.so runs each library's DT_INIT_ARRAY, an array of constructors: functions a library marks to run once at start-up, before any of its code is called. libc's run first, because the executable depends on it. Then the loader jumps to AT_ENTRY. That's your _start: a few instructions from crt1.o (a small piece of startup code gcc links into every program) that call __libc_start_main. The strings LD_DEBUG prints in section 5.2 are right here in the source:

C
#ifdef SHARED
  if (__builtin_expect (GLRO(dl_debug_mask) & DL_DEBUG_IMPCALLS, 0))
    GLRO(dl_debug_printf) ("\ninitialize program: %s\n\n", argv[0]);
 
  if (init != NULL)
    /* This is a legacy program which supplied its own init routine.  */
    (*init) (argc, argv, __environ MAIN_AUXVEC_PARAM);
  else
    /* This is a current program.  Use the dynamic segment to find
       constructors.  */
    call_init (argc, argv, __environ);      /* your DT_INIT_ARRAY */
  ...
  if (__glibc_unlikely (GLRO(dl_debug_mask) & DL_DEBUG_IMPCALLS))
    GLRO(dl_debug_printf) ("\ntransferring control: %s\n\n", argv[0]);
#endif
  __libc_start_call_main (main, argc, argv MAIN_AUXVEC_PARAM);
}

__libc_start_call_main calls main and passes its return value to exit. That init != NULL branch is also behind one of the most common deployment errors on Linux (section 6.2).

Hello made 98 symbol lookups before its first line ran, and a big C++ service can make tens of thousands. Many of those are for functions the program may never call. Looking all of them up in advance isn't required, and that's the next idea.

05Lazy binding and IFUNC

Relocations for function calls don't have to be done up front. With lazy binding, each one waits until the function is first called. And some symbols aren't looked up at all: for those, the loader runs a function to decide what the answer is.

5.1The PLT stub and its GOT slot

Calls to puts go through a small stub in the PLT (procedure linkage table), which has one stub per imported function. Each stub jumps through a GOT slot, and that slot initially points back into the loader. The first call resolves puts and patches the slot, and every call after that goes straight to libc.

?Is our hello lazily bound?

No, because of the FLAGS BIND_NOW in section 3.3. Ubuntu 24.04's gcc links with -z now by default, so every JUMP_SLOT is resolved at startup and the GOT becomes read-only. Making the GOT read-only after startup is called full RELRO (relocation read-only). It stops a memory-corruption bug, or an attacker, from overwriting a slot to redirect a call, but it only works if every slot is filled in first, so it needs eager binding. Watching a slot in gdb on a default build shows it filled before main. To see lazy binding at all, you have to ask for it with -Wl,-z,lazy.

Here's the PLT from the lazy build, objdump -d on aarch64:

Output
00000000000005d0 <.plt>:                    # PLT0: the path into the loader
 5d0:	stp	x16, x30, [sp, #-16]!
 5d4:	adrp	x16, 1f000
 5d8:	ldr	x17, [x16, #4088]            # GOT[2]: _dl_runtime_resolve
 5dc:	add	x16, x16, #0xff8
 5e0:	br	x17
 
0000000000000630 <puts@plt>:
 630:	adrp	x16, 20000                   # page of .got.plt
 634:	ldr	x17, [x16, #32]              # load the puts slot
 638:	add	x16, x16, #0x20              # x16 = &slot, for the resolver
 63c:	br	x17                          # jump wherever it points
 
0000000000000640 <main>:
 ...
 650:	bl	630 <puts@plt>

Each stub is four instructions. main calls the stub at 0x630, the stub loads whatever the puts slot holds, and jumps to it. The code at 0x5d0, labelled PLT0, is the way into the loader. Step through the first call and the second:

The first call to puts, lazily bound
mainhello's codeputs@pltfour-instruction stubGOTtable of addressesld.soPLT0 and _dl_fixuplibc.so.6the real putsbl puts@pltputs slotpoints to PLT0_dl_fixupputsbl
Step 1. main calls the stub, not puts itself: the compiler emitted bl 630 <puts@plt>. The puts slot in the GOT still holds the address of PLT0, the way into the loader.
1 / 6

5.2Watching one slot get patched

Stop on the bl in gdb (the standard debugger) and step over it, and you can watch the slot change:

Output
puts slot before the call:
0xaaaaaaac0020 <puts@got.plt>:	0x0000aaaaaaaa05d0      # PLT0, rebased
puts slot after the call:
0xaaaaaaac0020 <puts@got.plt>:	0x0000fffff7e619d0      # __GI__IO_puts in libc

Here's the function that did it:

C
_dl_fixup (struct link_map *l, ElfW(Word) reloc_arg)
{
  const ElfW(Sym) *const symtab = (const void *) D_PTR (l, l_info[DT_SYMTAB]);
  const char *strtab = (const void *) D_PTR (l, l_info[DT_STRTAB]);
  const uintptr_t pltgot = (uintptr_t) D_PTR (l, l_info[DT_PLTGOT]);
 
  const PLTREL *const reloc                 /* which .rela.plt entry? */
    = (const void *) (D_PTR (l, l_info[DT_JMPREL])
                      + reloc_offset (pltgot, reloc_arg));
  const ElfW(Sym) *sym = &symtab[ELFW(R_SYM) (reloc->r_info)];
  void *const rel_addr = (void *)(l->l_addr + reloc->r_offset);  /* the slot */
  ...
      result = _dl_lookup_symbol_x (strtab + sym->st_name, l, &sym, l->l_scope,
                                    version, ELF_RTYPE_CLASS_PLT, flags, NULL);
  ...
  if (sym != NULL
      && __builtin_expect (ELFW(ST_TYPE) (sym->st_info) == STT_GNU_IFUNC, 0))
    value = elf_ifunc_invoke (DL_FIXUP_VALUE_ADDR (value));   /* see 5.3 */
  ...
  /* Finally, fix up the plt itself.  */
  return elf_machine_fixup_plt (l, result, refsym, sym, reloc, rel_addr, value);
}

LD_DEBUG=bindings shows the order in which the bindings happen. On the lazy build, __libc_start_main binds after libc's constructors run, and puts binds after the loader prints transferring control, which means inside main:

Output
calling init: /lib/aarch64-linux-gnu/libc.so.6
binding file ./hello-lazy [0] to libc.so.6 [0]: normal symbol `__libc_start_main' [GLIBC_2.34]
initialize program: ./hello-lazy
transferring control: ./hello-lazy
binding file ./hello-lazy [0] to libc.so.6 [0]: normal symbol `puts' [GLIBC_2.17]

The full output, trimmed above, holds a surprise: hello-lazy appears to bind malloc, calloc, realloc and free, though it never calls them. That's ld.so itself. It runs on a tiny internal allocator until libc is relocated, then looks up the real malloc through the main program's scope. That's also how LD_PRELOAD gets to replace the loader's allocator.

5.3IFUNC: code that runs during relocation

memcpy in glibc isn't one function. readelf --dyn-syms on this libc marks it IFUNC, along with memmove, memset, memchr, strlen and gettimeofday. The fastest way to copy memory depends on the CPU, and libc is built before anyone knows which CPU it'll run on.

An IFUNC symbol points at a resolver. The loader calls it and uses the return value as the symbol's address. A resolver typically checks CPU features (that's what AT_HWCAP is for) and picks the SVE, SIMD or generic version of the function (SVE and SIMD are instruction sets that process several values at once).

When exactly does the resolver run? During relocation, inside the loader, while it's still patching tables, which means any library that ships an IFUNC gets to run its own code at that point. A test program makes it visible. It defines its own IFUNC sum, whose resolver records that it ran, and a constructor that checks. (The two versions of sum and the #include lines are left out.)

C
static int resolver_ran_before_main = 0;
static unsigned long hwcap_seen;
 
// Runs inside ld.so, while it processes this executable's relocations.
// On aarch64 the loader hands the resolver AT_HWCAP as its first argument.
static int (*resolve_sum(unsigned long hwcap))(const int *, int) {
    resolver_ran_before_main = 1;
    hwcap_seen = hwcap;
    return (hwcap & HWCAP_ASIMD) ? sum_simd : sum_generic;
}
int sum(const int *a, int n) __attribute__((ifunc("resolve_sum")));
 
__attribute__((constructor)) static void ctor(void) {
    printf("constructor: resolver already ran? %d\n", resolver_ran_before_main);
}

Running it:

Output
$ ./ifunc
constructor: resolver already ran? 1
main: sum=10, hwcap=0x40000000efb3ffff, ASIMD=yes

Our resolver ran before any constructor. It was triggered by a fourth kind of relocation, R_AARCH64_IRELATIVE, which means "call this resolver and store whatever it returns". Bit 62 in that hwcap is glibc's flag saying a second argument with more feature words follows.

Resolvers run this early by design, before RELRO makes the GOT read-only. That's convenient for the loader. Section 6.4 is about what happened when someone noticed it was convenient for an attacker too.

5.4What main can count on

Now we can say what the kernel and ld.so between them promise by the time control reaches main, and what they leave open:

SituationPromised?What it means
Every undefined symbol resolvesYes, with eager binding. With lazy binding, a missing function is only found at its first callWith -z now the process dies before main if a symbol is missing. With lazy binding it can die halfway through a request.
Constructors have runYesLibraries' before yours, dependencies before their dependents
The stack holds argc, argv, envp and the auxiliary vectorYesThe kernel's notes to user space, from section 4.2
The GOT is read-onlyOnly with full RELRONeeds -z now; a lazily bound program keeps a writable GOT
Which definition of malloc you getNoThe first object in search order that exports the name wins, and LD_PRELOAD puts itself first
The glibc on the machine is the one you built againstNoSection 6.2
How long any of this takesNoSection 7

The rows that aren't a plain yes are where programs get surprised, and the next section goes through those surprises one at a time.

06When the loader surprises you

Each of these cases has a cause you can read off the binary or its environment, and most of them play out before main, so your program never gets the chance to log what went wrong.

6.1'No such file or directory', for a file that's right there

Take a copy of hello and patch the path in its .interp section (the bytes that PT_INTERP points at) to name musl's loader, the path an Alpine binary would ask for. (musl is a small alternative C library, and Alpine Linux, popular in containers, uses it. Its loader lives at /lib/ld-musl-aarch64.so.1.) Then run it on Ubuntu.

Predict before you read on

The file exists and is executable. Its PT_INTERP names /lib/ld-musl-aarch64.so.1, which isn't installed. What does execve return?

Output
$ ls -l hello-badinterp
-rwxr-xr-x 1 root root 70312 Sep 25 13:26 hello-badinterp
$ ./hello-badinterp
bash: line 2: ./hello-badinterp: cannot execute: required file not found
$ strace ./hello-badinterp
execve("./hello-badinterp", ...) = -1 ENOENT (No such file or directory)

That's the open_exec failure from section 4.1. Bash 5.2 at least says "required file", and older shells just print "No such file or directory". That plain version is what Julia Evans saw when she ran into it with a glibc binary in an Alpine container.

That case was a loader that doesn't exist. A subtler case is a loader that does exist but is too old.

6.2version `GLIBC_2.34' not found

glibc gives its symbols version tags, and a program records the versions it was linked against. Look at what hello requires:

Output
$ objdump -T hello | grep GLIBC
0000000000000000      DF *UND*	0000000000000000 (GLIBC_2.34) __libc_start_main
0000000000000000  w   DF *UND*	0000000000000000 (GLIBC_2.17) __cxa_finalize
0000000000000000      DF *UND*	0000000000000000 (GLIBC_2.17) abort
0000000000000000      DF *UND*	0000000000000000 (GLIBC_2.17) puts

hello calls one function, puts, which has been in glibc forever, so the GLIBC_2.17 tag is no trouble. The GLIBC_2.34 tag comes from __libc_start_main, the function your _start calls, code you didn't write. So hello, built against glibc 2.34 or newer, refuses to start on anything older. In 2.34, __libc_start_main stopped using the init argument, and the glibc source says so:

Starting with glibc 2.34, the init parameter is always NULL. Older libcs are not prepared to handle that. The macro DEFINE_LIBC_START_MAIN_VERSION creates GLIBC_2.34 alias, so that newly linked binaries reflect that dependency. (csu/libc-start.c)

So anything built on Ubuntu 22.04 or newer won't start on a glibc 2.31 or 2.28 system, and the loader reports the error before main runs. Two working rules:

  • Build on the oldest system you ship to. glibc versioning is backward-compatible only: old binaries run on new glibc, never the reverse. Python's manylinux wheels are this rule turned into a standard (PEP 600).
  • Or don't depend on the host's libc at all: static linking, or a container that brings its own.

Both of those failures were accidents. The next case is the loader doing exactly what it was designed to do.

6.3Someone else's malloc: interposition and LD_PRELOAD

When the loader looks up a name, the first object in the search order that defines it wins. LD_PRELOAD puts an object at the front of that order, so its malloc beats libc's for every lookup in the process, libc's own included. This is symbol interposition, and it's why a profiler or a custom allocator can slip into a program that was never built for it.

?Why does a call inside a library go through the table too?

In a normal ELF shared library, a call from f1 to f2 inside the same library still goes through a table lookup. Some other object might define f2 first, and the rules say it wins. So the call needs a relocation like any other, even though both functions are in the same file.

Let's use the hook ourselves. A working malloc counter is 30 lines of C++. It finds the real malloc with dlsym(RTLD_NEXT, "malloc"), a call that looks up the next definition of a name after the current object in the search order, and it has to cope with dlsym itself calling malloc:

C++
// mallocount.cpp: count every malloc the program makes, print the total at exit.
#include <dlfcn.h>
#include <atomic>
#include <cstddef>
#include <cstdio>
#include <unistd.h>
 
using malloc_fn = void *(*)(size_t);
static malloc_fn real_malloc = nullptr;
static std::atomic<unsigned long> calls{0};
 
// dlsym can itself allocate. Serve those early requests from a static arena.
alignas(16) static char arena[4096];
static size_t arena_used = 0;
 
extern "C" void *malloc(size_t n) {
    if (!real_malloc) {
        static bool resolving = false;
        if (resolving) {                       // re-entered from inside dlsym
            void *p = arena + arena_used;
            arena_used += (n + 15) & ~size_t(15);
            return p;
        }
        resolving = true;
        real_malloc = reinterpret_cast<malloc_fn>(dlsym(RTLD_NEXT, "malloc"));
        resolving = false;
    }
    calls.fetch_add(1, std::memory_order_relaxed);
    return real_malloc(n);
}
 
__attribute__((destructor)) static void report() {
    char buf[64];
    int len = std::snprintf(buf, sizeof buf, "[mallocount] %lu calls\n",
                            calls.load(std::memory_order_relaxed));
    if (write(2, buf, len) < 0) {}              // not stdio: it may already be torn down
}

We build it as a shared library and preload it into four programs:

Output
$ g++ -std=c++20 -O2 -shared -fPIC -o mallocount.so mallocount.cpp
$ LD_PRELOAD=$PWD/mallocount.so ./allocs          # pushes 1,000 std::strings
[mallocount] 1013 calls
$ LD_PRELOAD=$PWD/mallocount.so python3 -c 'print(1)'
[mallocount] 3167 calls
$ LD_PRELOAD=$PWD/mallocount.so ./hello-static
hello
$ LD_PRELOAD=$PWD/mallocount.so ls /
$

Two of the runs counted allocations, and a thousand std::strings turn out to cost 1,013 calls. The last two are the interesting ones:

  • A static binary ignores LD_PRELOAD completely, since there's no loader to read it.
  • ls printed nothing. strace explains it: coreutils closes file descriptor 2, standard error, in its exit handler, and our destructor runs after that, so its write has nowhere to go.
Output
close(2)                                = 0
write(2, "[mallocount] 35 calls\n", 22) = -1 EBADF (Bad file descriptor)

On macOS the same shim counted zero calls, with or without DYLD_FORCE_FLAT_NAMESPACE. macOS binds each import to a named library (the two-level namespace), so defining malloc somewhere else doesn't change anything. You have to say what you're replacing in a __DATA,__interpose section. With that version, the same C++ program counted 1,038 calls.

So the loader lets a library run code before main (IFUNC resolvers, constructors) and lets it change which function a name means (interposition). In 2024 someone used both to reach into ssh servers.

6.4The xz backdoor: an IFUNC resolver doing too much

On 29 March 2024 Andres Freund posted to oss-security that xz-utils 5.6.0 and 5.6.1 contained a backdoor, now CVE-2024-3094. He'd noticed ssh logins on Debian sid taking a lot of CPU, plus valgrind errors, and measured a login going from 0.299 s to 0.807 s. Everything after that is loader mechanics. Some Linux distributions patch sshd, the server that handles ssh logins, so that it calls libsystemd (systemd's library for services to report that they're ready), and libsystemd depends on liblzma, the library xz-utils builds. That chain is what the attack used:

How the xz payload got from liblzma into sshd's RSA check
Libraries in sshdmapped by the loadersshd's GOTits table of addressesld.sorelocating, before mainlibsystemdservice notificationRSA_public_decryptpoints to libcryptoliblzma 5.6.xjust mappedaudit hookinstalled by liblzma
Step 1. A distribution-patched sshd links libsystemd. OpenSSH itself doesn't use liblzma. In the GOT sits the slot for RSA_public_decrypt, which sshd calls while checking a login's signature.
1 / 7

It only activated under narrow conditions, per Freund: x86-64 Linux, built by gcc and the GNU linker as a Debian or RPM package, argv[0] equal to /usr/sbin/sshd, TERM unset, LANG set, and LD_DEBUG and LD_PROFILE unset.

?Why didn't full RELRO stop it?

RELRO makes the GOT read-only after relocation. IFUNC resolvers run during relocation. So the attacker's code ran at a point when the tables it went after were still writable, by design.

Startup time was the clue that caught the backdoor, which is a good reason to know what normal startup costs.

07What startup costs

Everything above happens before main, so every process start pays for it. The timings in this section come from a small aarch64 Linux container. Each is a median over thousands of runs, where one run starts the program and waits for it to exit, and the rows were run interleaved so that background noise hits them all equally. The absolute numbers will differ on your machine; the gaps between rows are the part to carry away.

7.1Static, dynamic, C++, Python

ProgramMedianp10What's in it
hello, -static125 µs103 µs15 syscalls, no loader
hello, dynamic (-z now)154 µs133 µs32 syscalls, 98 symbol + 1,116 relative relocations
hello, dynamic, -z lazy152 µs132 µsSame, 5 slots deferred. No difference
hello with std::cout398 µs346 µslibstdc++, libm, libgcc_s. 1,794 symbol + 2,183 relative relocations
python3 -S -c pass3.4 ms3.2 mstimed separately with hyperfine, 5,000 runs
python3 -c pass4.5 ms4.0 mstimed separately with hyperfine, 5,000 runs

(The "p10" column is the tenth percentile: the time that 10% of runs beat, which is close to the best case without being a lucky outlier.)

Dynamic linking costs hello about 30 µs, maybe a little less. Swapping puts for std::cout costs another 240, most of it likely loading and relocating libstdc++. Python's 3–4 ms is mostly Python itself, the interpreter's own init and imports, and -S (skip site) takes a millisecond off.

Where the file lives matters too. With both binaries in a folder shared into the container from the host machine (a bind mount, which is slow to read from), a static hello looked slower than the dynamic one: 320 µs against 186 µs. The bigger static file paid more to be paged in. Copied to local disk, the p10s were 103 µs static and 132 µs dynamic.

?Does 30 µs matter?

For a long-lived server, probably not. But a build that runs the compiler 20,000 times pays it several times per file, since gcc execs cc1, as and ld. A shell script calling grep in a loop pays it every iteration.

Hello's 98 symbol relocations are cheap. What happens when a library has tens of thousands of them, and does lazy binding save us?

7.2Lazy versus now: a crossover

To make relocation cost visible, take a generated shared library holding a chain of about 20,000 functions: f0 calls f1, which calls f2, and so on down to f20000. With default visibility every one of those internal calls goes through the PLT, since any of them could be interposed. readelf -r counts 20,002 JUMP_SLOTs. With -fvisibility=hidden and only f0 exported: 2.

Predict before you read on

main calls f0, so all 20,000 functions run. Which build starts and finishes fastest?

Library build, what main callsBindingMedianSymbol relocations
default visibility, calls f19999 (2 functions run)lazy216 µs103
default visibility, calls f19999LD_BIND_NOW=1763 µs20,105
default visibility, calls f0 (all 20,000 run)lazy1,469 µs20,102, by the end
default visibility, calls f0LD_BIND_NOW=11,184 µs20,105
-fvisibility=hidden, calls f0either488 µs104
Eager binding, 20,000 extra symbols763 − 216 µs547 µs
Per symbol relocation547 µs ÷ 20,000≈ 27 ns
Lazy, all 20,000 called, over eager1,469 − 1,184 µs285 µs
Extra per lazily bound call285 µs ÷ 20,000≈ 14 ns
Hidden visibility vs eager, same work−696 µs, 2.4× faster

So which wins? Lazy binding wins by 3.5× when you call a handful of what you link, and loses by 24% when you call everything. Most real programs call a small fraction of libc and a large fraction of their own code, which argues for lazy. Full RELRO needs eager binding, though, and that's why distributions pay anyway.

?Why does a lazy first call cost 14 ns more than eager binding?

These numbers don't settle it. _dl_runtime_resolve stores and reloads 208 bytes of registers on every first call, which should cost a few nanoseconds, and the lookup is the same one eager binding does. A likely candidate is page faults: eager binding writes the GOT in one sequential pass, while lazy binding writes one slot at a time, scattered through the run. Running perf stat -e page-faults,dTLB-load-misses on both variants would separate the causes, but it needs hardware performance counters, which containers often hide.

Large programs with huge numbers of relocations were a real problem long before anyone timed them this way, and people have built several fixes for it.

7.3The history of trying to make this cheaper

Big C++ programs, office suites and browsers especially, are where this hurt, and the software on your machine probably carries the result:

FixWhat it didWhere it went
prelink (Jakub Jelinek, Red Hat, paper)Computed relocations ahead of time for a fixed library layoutFedora and RHEL ran it by default. Jelinek reported "an order of magnitude difference" in relocation time for OpenOffice.org Writer (LWN, 2009). It's at odds with ASLR, since the whole point is that addresses stay put.
Firefox's elfhack (Mike Hommey, 2010)Rewrote libxul's relative relocations into a compact form to cut startup I/O (glandium.org)Retired in 2023 in favour of RELR (glandium.org)
RELR (DT_RELR)Standardised the compact relative-relocation formatChrome OS carried patches from 2018, Android and Fuchsia adopted it, and it landed in glibc 2.36 (MaskRay). Linkers emit it with -z pack-relative-relocs.

The reason so many teams bothered is size as well as time. Each ordinary relative relocation takes 24 bytes in the file, and MaskRay found relative relocations made up 7.9% of total file size across Arch Linux's /usr/bin.

All of these keep the loader and make it cheaper. Another answer is to have no loader at all.

08Static by default: Go and Alpine

Every cost and failure so far traces back to one decision, resolving addresses at run time. Some ecosystems decided otherwise, or made a different choice of C library, and you'll meet both in containers.

8.1Go: static by default

Go sidesteps GLIBC_x.y errors and missing libraries by linking statically. With cgo off (cgo is Go's bridge to C libraries), a Go binary has no PT_INTERP, so it runs even in a scratch container (an empty image with no files but yours) that has no libc at all.

One catch is that a static Go binary can't use glibc's DNS resolver, the code that turns host names into network addresses (a different "resolver" from the IFUNC kind). The net package prefers its own pure-Go resolver on Unix, and falls back to glibc's through cgo when the system's name-lookup configuration (/etc/nsswitch.conf, /etc/resolv.conf) asks for features the Go one doesn't implement. GODEBUG=netdns=go or =cgo forces one; netdns=1 logs the choice. If a Go service resolves names differently from curl on the same box, check which resolver it's using.

8.2Alpine and musl

A glibc binary on Alpine fails with the ENOENT from section 6.1, because Alpine's loader is musl's. And a musl binary behaves differently from glibc in ways musl documents: no lazy binding at all, dlclose as a no-op, and a resolver that didn't support DNS over TCP until 1.2.4 in May 2023. Before that, a truncated UDP answer just failed, so a name that resolved fine on your laptop could fail under a big DNS answer.

With the choices laid out, here's how to see what a given binary is doing and how to decide.

09Seeing inside the loader

9.1The commands

Each question this chapter raised has a command that answers it.

Shell
# What will run first, and what does this need? (sections 3 and 4)
readelf -lW ./app | grep -A1 INTERP
readelf -dW ./app | grep -E 'NEEDED|RUNPATH|RPATH|FLAGS'
objdump -T ./app | grep -o 'GLIBC_[0-9.]*' | sort -V | tail -1   # minimum glibc (section 6.2)
 
# RELRO status: GNU_RELRO present + BIND_NOW = full RELRO (section 5.1)
readelf -lW ./app | grep GNU_RELRO
readelf -dW ./app | grep -E 'BIND_NOW|FLAGS_1'
 
# The loader's own diagnostics (sections 1.2, 4.3 and 5.2)
LD_DEBUG=libs ./app          # search paths, which file each soname became
LD_DEBUG=bindings ./app      # every symbol, who asked, who answered
LD_DEBUG=statistics ./app    # relocation counts
LD_DEBUG=help ./app          # the list
LD_SHOW_AUXV=1 ./app         # what the kernel handed over (section 4.2)
 
# Startup syscalls (section 4.3)
strace -f -e trace=openat,mmap,mprotect ./app

9.2Rules that hold up

  1. Build shared libraries with -fvisibility=hidden and export an explicit API. Smaller file, fewer lookups, faster calls. It was the biggest single win in section 7.2.
  2. Keep full RELRO (-z relro -z now). Ubuntu's default already does. It costs eager binding, and in exchange an attacker can't overwrite GOT entries after startup.
  3. Link RELR with -z pack-relative-relocs if your toolchain and target glibc (2.36 or newer) support it.
  4. Count your libraries. Each NEEDED is an openat, several mmaps, relocations and its constructors. ldd on a large C++ service can easily list dozens.
  5. Go static for tools you exec a lot, if you can live with the trade-offs in the next table.

9.3What you trade for what

ChoiceYou getYou pay
Dynamic linkingShared pages, security fixes without rebuilds, LD_PRELOAD~30 µs per exec for hello, far more for C++; GLIBC_x.y errors
Static linkingOne file, no loader, no version skewNo LD_PRELOAD, rebuild for every libc CVE, and glibc's NSS/dlopen features break
Lazy binding3.5× faster startup when few symbols are calledWritable GOT for the process lifetime; a missing symbol is a crash mid-run
Full RELRO (-z now)Read-only GOT after startupEvery symbol bound up front, used or not
-fvisibility=hiddenFewer relocations, direct calls, smaller fileYou must mark the API explicitly; nobody can interpose internals
musl (Alpine)Small, simple, static-friendlyNo lazy binding, dlclose is a no-op, DNS over TCP only since 1.2.4

9.4Symptom, cause, fix

SymptomLikely causeFix
"No such file or directory" (or "required file not found") for a file that existsPT_INTERP names a loader that isn't installed, often a glibc binary on Alpinereadelf -l ./bin | grep interpreter; build for the target's libc
version `GLIBC_2.34' not foundBuilt against a newer glibc than the target hasBuild on the oldest target, use a manylinux-style image, or link statically
cannot open shared object fileThe loader can't find a NEEDED sonamereadelf -d for NEEDED and RUNPATH; LD_DEBUG=libs for the search
LD_PRELOAD has no effectThe binary is static, or it's macOS's two-level namespaceCheck for PT_INTERP; on macOS use __DATA,__interpose
A preload shim crashes or hangs at startupdlsym allocating and re-entering your mallocServe re-entrant calls from a static arena
Slow start for a big C++ serviceTens of thousands of symbol relocations-fvisibility=hidden, fewer libraries, RELR
A Go service resolves names differently from curlThe pure-Go resolverGODEBUG=netdns=1 to see which, =go or =cgo to force one

10Summary

  1. A compiled program calls functions whose code isn't in its file, so someone has to supply their addresses at run time. Static linking copies the code in; dynamic linking shares one copy and fills in addresses at start-up.
  2. The CPU starts in ld.so, not your program. The kernel maps both, reads PT_INTERP, and hands your entry point over as AT_ENTRY.
  3. An ELF file describes itself twice. Program headers are for loading; section headers are for tools, and the kernel never reads them.
  4. A relocation is "write this address here". Relative ones are cheap; symbol ones need a lookup through every loaded object.
  5. GNU hash makes "not here" cheap. A Bloom filter rejects most objects with one load and two shifts.
  6. Ubuntu binds everything at startup by default. -z now gives full RELRO and a read-only GOT; lazy binding needs -z lazy.
  7. IFUNC resolvers run inside the loader, before RELRO. That's the door the xz backdoor used to reach sshd.
  8. "No such file or directory" usually means the interpreter. The kernel returns ENOENT for the loader that isn't there.
  9. Even hello needs GLIBC_2.34 if you build it on a recent distribution, so build on the oldest system you ship to.
  10. LD_PRELOAD wins every lookup, and static binaries ignore it.
  11. Hidden visibility beats both binding modes. 488 µs against 1,184 for the same 20,000 calls, and a library half the size.

11Build this

Write an LD_AUDIT library that times every binding in a real program.

glibc's rtld-audit(7) interface calls your la_objopen for every library loaded and la_symbind64 for every symbol bound. That's more or less the same machinery the xz payload hooked.

  • Implement la_version, la_objopen and la_symbind64, and log object name, symbol name and a clock_gettime timestamp for each.
  • Run it on python3 -c pass and on a C++ service you own. Count bindings per library.
  • Then rebuild one of your own libraries with -fvisibility=hidden and watch its bindings disappear from the log.

Your first run will probably show you a library you didn't know was loaded. On Debian and Ubuntu, liblzma in sshd was exactly that.

12Interview questions

beginnerWhat's the difference between program headers and section headers?›

Program headers describe segments: what to map, where, with which permissions. That's all the kernel and the dynamic loader read. Section headers describe the file for tools like the linker, objdump and gdb: .text, .dynsym, .rela.plt. Linux never looks at them, so a binary can run with its section header table stripped.

beginnerA binary exists and is executable, and running it says 'No such file or directory'. Why?›

Almost always the interpreter named in PT_INTERP doesn't exist, typically a glibc binary in an Alpine container asking for /lib/ld-linux-*.so. Linux opens that path inside load_elf_binary and returns its ENOENT from execve. readelf -l ./bin | grep interpreter shows what it wants.

intermediateWalk through the first call to puts in a lazily bound program.›

main calls puts@plt. That stub loads the puts slot from .got.plt. It still points at PLT0, and jumps there. PLT0 pushes the slot address and jumps to _dl_runtime_resolve. That saves the argument registers and calls _dl_fixup. That finds the .rela.plt entry, looks puts up through the GNU hash tables of each loaded object, writes the address into the slot, and returns it. The trampoline restores registers and jumps to puts. Later calls skip all of it.

intermediateWhy does a program built on Ubuntu 24.04 fail with GLIBC_2.34 not found on an older box?›

glibc versions its symbols, and a program records the version it linked against. Even a hello world references __libc_start_main@GLIBC_2.34, from crt1.o, because 2.34 changed how that function handles constructors. Older glibc doesn't have that version node, so the loader refuses before main. Build on the oldest target, use a manylinux-style build image, or link statically.

intermediateWhat does full RELRO protect, and what does it cost?›

It makes the GOT and other relocated data read-only once the loader finishes, so a memory-corruption bug can't redirect a function pointer in the GOT. It requires binding every symbol at startup (-z now). In the 20,000-symbol test that cost 547 µs against lazy binding when only two functions were called.

deepHow did the xz backdoor get code running inside sshd, and why didn't RELRO help?›

Distro-patched sshd links libsystemd, which links liblzma. That malicious liblzma replaced its IFUNC resolvers. The loader calls those during relocation, before main and before RELRO is applied. From there it installed a dynamic-linker audit hook and redirected RSA_public_decrypt when that symbol was bound. RELRO locks tables after relocation, and the attack happened during it.

deepWhy does -fvisibility=hidden make a shared library start faster?›

Default visibility means every exported function can be interposed by an earlier object, so even internal calls go through the PLT and need a JUMP_SLOT relocation. Hidden symbols can't be interposed, so the static linker emits direct branches and drops them from .dynsym. In the test that took a library from 20,002 jump slots to 2, from 3.7 MB to 2.0 MB, and cut startup from 1,184 µs to 488.

deepYour LD_PRELOAD malloc shim deadlocks or crashes at startup. What's the likely cause?›

Probably recursion during bootstrap. Your shim's malloc calls dlsym(RTLD_NEXT, "malloc") to find the real one, and dlsym itself can allocate, and that re-enters the shim before the pointer is set. The usual fix is a static arena for re-entrant calls during resolution, or a glibc-specific entry point. Printing from a destructor has its own trap: some programs close stderr before destructors run.

13Go deeper

check yourself
Where does the CPU start executing after execve of a dynamic PIE?›

In ld.so, at the interpreter's entry point. The program's own entry goes on the stack as AT_ENTRY and ld.so jumps there after relocation.

Why did LD_PRELOAD do nothing to the static binary?›

There's no dynamic loader to read the variable. Static binaries have no PT_INTERP and resolve everything at link time.

Lazy binding or BIND_NOW: which starts faster?›

It depends on how much you call. Lazy was 3.5× faster when two of 20,000 functions ran, and 24% slower when all of them did.

Your hello world was built on Ubuntu 24.04 with gcc -O2. Is it lazily bound?›

No. Ubuntu's toolchain links with -z now by default, so FLAGS shows BIND_NOW and the GOT is read-only before main.

Operating Systems: Three Easy Pieces, Interlude: Process API (chapter 5)

How fork() and exec() start a new program, the step just before everything in this chapter. It doesn't cover ELF or the loader, so it pairs with sections 4 and 5 here. Free online.

Ulrich Drepper — How To Write Shared Libraries

The paper on relocation cost, visibility and GNU hash, from glibc's former maintainer. Old and still correct on every mechanism in section 3.

fs/binfmt_elf.c at v6.10

load_elf_binary and create_elf_tables. About 500 lines, and all of sections 4.1 and 4.2.

glibc elf/ at 2.39

dl-runtime.c, dl-lookup.c and the aarch64 dl-trampoline.S. Read _dl_fixup first.

MaskRay — Relative relocations and RELR

The history from elfhack to DT_RELR, with size measurements across a whole distribution.

04 — Virtual Memory & Page Tables

Chapter 04. Every mmap in the strace above creates a mapping whose pages fault in on first touch. That's part of why a binary on a bind mount started so much slower than the same binary on local disk.

07 — Syscalls & the Kernel Boundary

Chapter 07. execve is the heaviest syscall most programs make, and AT_SYSINFO_EHDR in the auxv is where the vDSO from that chapter gets handed over.

05 — Allocators

Chapter 05. What you're replacing when an LD_PRELOAD of jemalloc or tcmalloc interposes malloc.

06 — Processes & Scheduling

Chapter 06. The fork half of fork-and-exec, and what the new process inherits before the loader runs.