KnowSys
▣Build it yourself

Build Docker, from scratch

A container runtime in a few hundred lines: namespaces, a root filesystem, cgroup limits, layered images and a network, until docker run stops being magic.

Intermediate⏱ 4–6 weekendsGo · Rust · C

docker run alpine sh feels like starting a tiny virtual machine. It isn't. It's one ordinary Linux process that the kernel has been asked to lie to in a few specific ways, plus a cgroup that bills it for what it uses.

Building your own runtime is the fastest way to see that. By the end you'll have a program that pulls a real image from Docker Hub, unpacks its layers, starts a process that thinks it's PID 1 on its own machine, with its own filesystem, hostname and network, and kills it if it uses more than its memory limit. It's probably a few hundred lines, and every line maps to a kernel feature you can name.

01Why build this

Almost everything you deploy now runs in a container, and most engineers treat the runtime as a black box. Building one changes how you debug production:

  • OOM kills stop being mysterious. You'll have written the memory.max line that triggers them.
  • "Works on my machine, not in the pod" gets easier. You'll know exactly which parts of the environment a container replaces and which it shares, like the kernel.
  • Security reviews make sense. You'll know why a container isn't a VM, and which of capabilities, seccomp and user namespaces closes which hole.
  • Kubernetes gets thinner. Under the CRI, containerd and runc do what your program will do.

It's also probably one of the best-scoped systems projects there is. Each milestone runs on its own and is useful, so you can stop at any point with something that works.

02What you're building

The finished runtime does these steps for every run:

What your runtime does on `run alpine sh`
●
⬇
Image
pull + unpack
▦
Rootfs
overlayfs
◌
Namespaces
clone()
▤
cgroup
limits
⇄
Network
veth + bridge
▶
exec
PID 1
Step 1. Fetch the image manifest from the registry and download its layers: gzipped tarballs, each one a set of file changes.
1 / 6

?Why is there no "container" object in the kernel?

Because Linux never added one. Namespaces, cgroups and mounts are separate features, added over about a decade, that happen to compose. A "container" is just the name for a process with all of them applied. That's why you can build one piece at a time.

03Before you start

You needWhyWhere to get it
A Linux machine you can be root onNamespaces and cgroups are Linux kernel featuresA VM, or docker run --privileged on a Mac
Kernel 5.x or newer, cgroup v2The limits API used belowAny current distro
Comfort with fork, exec and file descriptorsYou'll call them directlyChapter 07
One systems languageDirect syscallsGo is easiest, Rust is safest, C shows the most

04The roadmap

Eight milestones. Each one is usable on its own, and each introduces one kernel idea.

1

A process that thinks it's alone

1 evening

Start the child with clone() (in Go, SysProcAttr.Cloneflags) and ask for new UTS and PID namespaces. Set a hostname inside. The child now has its own hostname and sees itself as PID 1.

Run ps inside and you'll still see every host process. That's the first surprise, and the next milestone fixes it: ps reads /proc, and you haven't given the child its own yet.

You’ll learnclone()CLONE_NEWPIDCLONE_NEWUTSPID 1
Done when: hostname inside differs from the host, and echo $$ prints 1.
2

Its own root filesystem

1 weekend

Download Alpine's "minirootfs" tarball and unpack it into a directory. Add CLONE_NEWNS, make every mount private so nothing propagates back to the host, then pivot_root into the directory and mount a fresh /proc.

Use pivot_root, not chroot. A chroot can be escaped by a root process; pivot_root swaps the real root and lets you unmount the old one completely.

You’ll learnmount namespacepivot_rootMS_PRIVATE/proc
Done when: ls / shows Alpine's tree, ps shows only your process, and the host's mount table is untouched.
3

Limits with cgroups v2

1 evening

Create a directory under /sys/fs/cgroup, write 64M to memory.max, 50000 100000 to cpu.max (half a CPU), and 64 to pids.max. Write the child's PID into cgroup.procs before it execs.

Then break it on purpose. Allocate past the limit and watch the kernel kill it. Run a fork bomb and watch it stop at 64 processes. That's essentially all a Kubernetes resource limit is.

You’ll learncgroupfsmemory.maxcpu.maxpids.maxOOM kills
Done when: A memory hog inside is killed at your limit, and memory.events shows oom_kill 1.
4

Layered images with overlayfs

1 weekend

Keep the unpacked image read-only as lowerdir, and give each container its own empty upperdir and workdir. Mount the overlay and pivot into that.

Write to a file that came from the image, and overlayfs copies it up into the upper layer first. Delete one, and it leaves a whiteout marker. You'll need both ideas in milestone 6.

You’ll learnoverlayfslowerdir / upperdircopy-upwhiteouts
Done when: Two containers share one read-only image; changes in one never appear in the other or in the image.
5

A network of its own

1 weekend

Add CLONE_NEWNET, and the child has only a loopback interface. Create a veth pair (two interfaces joined like a cable), move one end into the child's namespace, and attach the other to a bridge on the host. Give each end an address, add a default route inside, and set up NAT on the host so traffic can leave.

Copy a resolv.conf in, or DNS won't work. That's probably the most common "networking is broken" bug at this stage.

You’ll learnnetwork namespaceveth pairsbridgesNATresolv.conf
Done when: The container can ping another container by IP and fetch a web page from the internet.
6

Pull real images

1–2 weekends

Talk to a registry over HTTPS. Get a token, fetch the manifest for a tag (choosing your CPU architecture from the index), then download each layer by its sha256 digest. Every blob is content-addressed, so verify the digest as you download.

Unpack the layers in order, applying whiteouts, and you have a lowerdir for milestone 4. This is also where you learn why images are cached by digest, not by name.

You’ll learnOCI distribution APIimage manifestscontent addressinglayer tarballs
Done when: run nginx pulls from Docker Hub, unpacks every layer in order, and serves a page.
7

Make it safe

1 weekend

Root inside your container is still root on the host kernel. Close that down in layers:

  • Drop capabilities to the short list Docker uses, so "root" can't mount, load modules or change the time.
  • Set no_new_privs, so setuid binaries can't grant anything back.
  • Add a seccomp filter that refuses dangerous syscalls.
  • Try user namespaces last, mapping container root to an unprivileged host user. That's "rootless" containers.
You’ll learncapabilitiesseccompuser namespacesno_new_privs
Done when: Root inside the container can't mount filesystems, load kernel modules, or change the host clock.
8

Lifecycle: PID 1, signals and exec

1 weekend

PID 1 has duties no other process has: it must reap orphaned children and it only receives signals it has handlers for. Either make your runtime inject a tiny init (the way docker run --init adds tini), or reap and forward signals yourself.

Then add exec: open the running container's namespace files in /proc/<pid>/ns/ and setns() into each one before running a new command. That's exactly what docker exec and kubectl exec do.

You’ll learnzombie reapingsignal forwardingsetns()exit codes
Done when: stop ends the container cleanly, no zombies remain, and exec opens a shell in a running container.

05Traps that catch everyone

SymptomCauseFix
Mounts made inside the container appear on the hostMount propagation is shared by defaultRemount / as MS_PRIVATE or MS_SLAVE before anything else
ps inside shows host processes/proc is still the host'sMount a new proc after pivot_root
The host loses its networkYour script moved the host's real interface into the namespaceAlways create and move a fresh veth end; never touch eth0
The memory limit does nothingYou tested from nsenter, which doesn't join the cgroup, or the kernel uses cgroup v1Put the PID into cgroup.procs; check stat -fc %T /sys/fs/cgroup says cgroup2fs
Deleted files reappear from the imageWhiteout files were ignored while unpackingHandle .wh. entries and opaque directories
docker stop takes 10 secondsYour PID 1 ignores SIGTERMForward signals, or run a minimal init
DNS fails, ping by IP worksNo resolv.conf in the rootfsWrite one in, or bind-mount the host's

06Stretch goals

  • Speak the OCI runtime spec. Read a config.json bundle like runc does, and containerd can run your runtime in place of runc.
  • Rootless by default. Run the runtime as a normal user with user namespaces and newuidmap.
  • Checkpoint and restore. Freeze a container with the cgroup freezer and look at how CRIU serialises a running process.
  • A tiny orchestrator. Run several containers from a YAML file with a shared network and restart policies: a miniature docker compose.

07References worth your time

Liz Rice, Containers From Scratch

A talk that builds a container live in Go in about 20 minutes. The best first hour you can spend on this project.

Linux containers in 500 lines of code

Lizzie Dixon's annotated C runtime, with capabilities, seccomp and user namespaces done properly.

bocker

Docker's core ideas in about 100 lines of bash. Useful for seeing how small the essential logic is.

OCI runtime and image specs

What runc and every registry actually implement: the bundle format, the manifest, and the layer rules.

man7.org: namespaces(7), cgroups(7)

The authoritative description of every flag and file you'll touch.

runc and crun source

Production runtimes in Go and C. Read them after milestone 8, when you'll recognise every piece.

Chapters that back this project

Next project⛁ a database storage engine→