Every program you've written has run with something underneath it, catching
its page faults, deciding when it gets the CPU and turning write into bytes on
a device. Writing a kernel means being the thing underneath. Nobody catches
your mistakes; when you get one wrong, the virtual machine resets and you get
no error message at all.
Plenty of hobby kernels stop at "a worse Unix", and that's where they lose steam. This one ends differently. Its first seven milestones build the parts every kernel needs, and the last one asks you to pick a superpower: a unikernel that boots straight into one application, a real-time scheduler that guarantees deadlines, or a microkernel that moves drivers out into user space and connects them with message passing. That choice is what makes the project yours, and it's the part you'll end up explaining to people.
01Why build this
Most of what's slow, surprising or dangerous in production happens at the kernel boundary. Having written one, you'll read those problems differently:
- Page faults, TLB misses and huge pages get concrete. You'll have written the page tables and flushed the TLB by hand. Chapter 04 becomes a description of your own code.
- Syscall overhead stops being folklore. You'll know exactly which registers get saved, which stack gets switched and why a mode change costs something.
- Scheduling decisions make sense. Preemption, time slices and priority inversion are things you'll have caused and debugged.
- Allocator behaviour becomes intuitive. Your first heap will fragment, and you'll see why.
- Your superpower teaches a design trade-off. Unikernels, real-time kernels and microkernels each give up something a general-purpose kernel keeps, and you'll know what and why.
It's also the project where "it works" feels best. There's nothing between your code and the CPU except QEMU.
02What you're building
From power-on to a running user program, your kernel does this:
?Why do the trap handlers come before memory management?
Because without them every bug looks the same. If the CPU hits a fault and there's no handler, it faults again trying to report that, and a third fault resets the machine. On x86 that's a triple fault, and all you see is QEMU rebooting. With handlers in place, a bad page table prints the faulting address and you can fix it in minutes. Install them early and everything after gets easier.
03Before you start
| You need | Why | Where to get it |
|---|---|---|
| QEMU | Reboots in a second, and it can tell you why a fault happened | Your package manager: qemu-system-x86_64 or qemu-system-riscv64 |
| A freestanding toolchain | There's no OS or C library under you | Rust with a bare-metal target, Clang, or a cross GCC like riscv64-unknown-elf-gcc |
| A bootloader or firmware | Gets you into 64-bit mode with a memory map | Limine or GRUB with Multiboot2 on x86-64; OpenSBI ships with QEMU for RISC-V |
| GDB | Single-step your kernel | Start QEMU with -s -S, then target remote :1234 |
| The architecture manual | The final word on every register | Intel SDM Volume 3, or the RISC-V privileged specification |
| Virtual memory and syscall basics | You're about to implement both | Chapter 04, Chapter 07 |
04The roadmap
Eight milestones. Numbers 1 to 7 are the kernel every design needs; number 8 is where you make it unusual.
Boot and say hello
1 weekendBuild a binary with no standard library and a linker script that puts your entry point where the bootloader expects it. Write a few lines of assembly to set up a stack and call your main function.
Then print. On x86-64, write bytes to the COM1 serial port; on RISC-V's QEMU
virt machine, write to the 16550 UART. Pass -serial stdio or -nographic
so the output lands in your terminal. You'll use that print function for the
rest of the project, so make it formatted.
Catch faults
1 weekendInstall a handler table. On x86-64 that's the IDT, with small assembly
stubs that save registers into a trap frame before calling your code. On
RISC-V it's one entry point in stvec, and scause tells you what happened.
Handle a breakpoint first, since it's harmless and returns. Then page faults,
printing the address from CR2 or stval. On x86-64, give the double-fault
handler its own stack through the TSS; otherwise a kernel stack overflow faults
again inside the handler and you're back to silent resets.
Interrupts: a timer and a keyboard
1 weekendEnable hardware interrupts. On x86-64, remap the legacy PIC so its vectors don't collide with CPU exceptions, or set up the APIC. On RISC-V, ask OpenSBI to arm the timer and use the PLIC for the UART. Count timer ticks, and echo keyboard or serial input.
Two things will bite here. Forget to acknowledge an interrupt and it fires once, then never again. And if a handler takes a lock the interrupted code already holds, you deadlock; disable interrupts while holding any lock a handler also takes.
Paging and physical memory
1–2 weekendsParse the bootloader's memory map and build a frame allocator that hands
out free 4 KiB physical pages. A bitmap or a free list is fine. Then write
map(virtual, physical, flags) for four-level x86-64 tables or RISC-V's Sv39.
Switch to your own tables carefully: the code doing the switch, its stack and
the serial port all have to be mapped in the new tables, or the next
instruction fetch faults. After changing a mapping, flush it with invlpg or
sfence.vma. Stale TLB entries cause bugs that look random.
A kernel heap
1 weekendMap a region for the heap and start with a bump allocator. It's a few lines and can't free anything, which you'll notice soon. Replace it with a free-list allocator that splits and merges blocks, then try fixed size classes for small objects.
Hook it into your language so ordinary collections work: GlobalAlloc in Rust,
malloc and free in C. Run a stress test that allocates in a random pattern
and watch fragmentation happen. Chapter 05 explains what
you're seeing.
Threads and a preemptive scheduler
1–2 weekendsGive each thread its own kernel stack and a saved-register area. A context switch is a short assembly routine: save the callee-saved registers and stack pointer of one thread, load another's, return. Seeing it work the first time is one of the best moments in the project.
Then make it preemptive. On each timer tick, the handler asks the scheduler for the next runnable thread and switches to it. Start with round robin. Add sleeping and a way for a thread to block waiting on an event.
User mode and system calls
2 weekendsGive each process its own page table, with the kernel mapped into all of them
but marked supervisor-only. Load a small ELF program into the user half,
and drop to user mode with sysret or sret.
Programs come back in through the syscall instruction on x86-64 or ecall
on RISC-V. Implement write, exit and yield, and check every pointer a
program passes you: it must be inside that process's user memory. A kernel
that trusts user pointers is the textbook privilege escalation.
Pick your superpower
2–4 weekendsChoose one, and let it reshape the kernel:
- A unikernel. Link a single application, such as a key-value server over the serial port or virtio-net, straight into the kernel. One address space, no user/kernel switch, no scheduler to speak of. Measure time from power-on to first response, and write down exactly what isolation you gave up.
- A real-time kernel. Tasks declare a period, a worst-case execution time and a deadline. Schedule with earliest deadline first or rate-monotonic priorities, refuse tasks that would push utilisation past what the algorithm can guarantee, and log every miss and its jitter. Add priority inheritance so a low-priority task holding a lock can't stall a high-priority one.
- A microkernel. Move the serial driver and a tiny filesystem into user processes. The kernel keeps only address spaces, threads and synchronous message-passing IPC: send, receive and reply on endpoints. Measure an IPC round trip; making it fast is the core problem Liedtke's L4 work attacked.
Whichever you pick, write down the numbers. A kernel that can say what it's
good at, and prove it, is worth far more than one that runs a copy of ls.
05Traps that catch everyone
| Symptom | Cause | Fix |
|---|---|---|
| QEMU reboots in a loop | A triple fault: a fault while handling a fault, often from a bad stack or missing handler | Run with -d int,cpu_reset -no-reboot; give double faults their own stack |
| An interrupt fires once, then never again | No end-of-interrupt was sent to the controller | Send EOI to the PIC or APIC, or complete the claim on the PLIC |
| Random corruption once interrupts are on | Your stub doesn't save every register, or the stack isn't 16-byte aligned | Save the full trap frame in assembly; align before calling into your language |
| A crash right after switching page tables | The running code, stack or UART isn't mapped in the new tables | Map the kernel and its stack in every address space before switching |
| Works in debug, crashes in release on x86-64 | The compiler uses the red zone or SSE registers your handlers don't save | Build the kernel with the red zone and SSE disabled |
| A mapping change seems to be ignored | Stale TLB entry | invlpg or sfence.vma after every change |
| The system freezes in an interrupt handler | A handler waits on a lock the interrupted code holds | Disable interrupts while holding locks shared with handlers |
06Stretch goals
- Multiple cores. Start the other CPUs, give each its own run queue and per-CPU data, and find every place a spinlock was quietly assumed away.
- A block device and a filesystem. Write a virtio-blk driver and a small filesystem on top; xv6's log-based one is a good model.
- Networking. A virtio-net driver plus a minimal IP, UDP and TCP stack, so your unikernel or microkernel can talk to the host.
- Real hardware. Boot on an old PC from USB or on a RISC-V board. Expect timing and firmware differences QEMU hid from you.
- A second superpower. Combine two, such as a microkernel with a real-time scheduler, and find where their goals conflict.
07References worth your time
A blog series that takes an x86-64 kernel from a bare binary through interrupts, paging, heap allocation and async tasks. Milestones 1 to 5 follow roughly the same order.
The companion to MIT's teaching kernel, a small Unix for RISC-V. Its chapters on traps, page tables and scheduling are short and exact.
The community reference for every piece of PC hardware and every boot protocol, plus a long list of mistakes other people already made.
A free textbook on virtualisation, concurrency and persistence. Read the scheduling and memory chapters alongside milestones 4 to 6.
The 1995 SOSP paper arguing that microkernels were slow because of how they were built, not what they were. Required reading for the microkernel option.
The 2013 ASPLOS paper behind MirageOS, which set out the case for single-application kernels.
The 1973 paper that introduced the rate-monotonic utilisation bound and showed EDF's optimality. Everything in the real-time option builds on it.