A browser looks like the most complicated program on your machine, and the production ones are. But the core of one is a pipeline you can hold in your head: fetch some bytes, turn them into a tree, decide what each node looks like, decide where each node goes, and draw it. Chrome and Firefox are that pipeline plus twenty-five years of edge cases and speed.
You're going to build the pipeline. By the end you'll have a program that opens a real HTTPS page, parses the HTML into a DOM, applies a stylesheet, lays out blocks and wrapped lines of text, paints the result into a window, scrolls, and navigates when you click a link. If you have the energy after that, you'll plug in a JavaScript engine and watch a script change the page.
01Why build this
The browser is the runtime for most of the software people use, and most engineers who write for it have never seen its insides. Building one changes a few things:
- Front-end performance advice starts to make sense. "Avoid layout thrashing" and "animate transforms, not widths" are obvious once you've written the code that has to redo layout.
- HTTP stops being a library call. You'll write request lines and headers by hand, handle chunked bodies, and find out what keep-alive really costs.
- You learn to read specs. The HTML and CSS specs are unusually precise documents, and being able to find an answer in one is a skill that transfers to RFCs and every other standard.
- Trees, caches and invalidation everywhere. A browser is a lesson in keeping several derived data structures in sync with a source that keeps changing. So is most backend work.
It also has an unusually good test suite: the entire web. Point it at a page you like and see what breaks.
02What you're building
Loading a page runs these stages, each producing a new data structure from the previous one:
GET with a Host header, and read back the status, headers and body.?Why keep four separate trees and lists?
Because they change at different rates. Scrolling only needs a new raster. Changing a colour needs new style and paint but not new layout. Changing a width needs layout again. If the DOM, styles, layout and display list were one structure, every change would redo everything. Real engines spend enormous effort on exactly this invalidation problem, and keeping your stages separate lets you see where it comes from.
03Before you start
| You need | Why | Where to get it |
|---|---|---|
| A language you're fast in | This is a lot of code; fighting the language slows you down | Python is fine and is what browser.engineering uses; Rust or C++ if you want speed |
| A window with 2D text and shape drawing | Milestone 2 onwards draws to the screen | Tkinter, SDL with a font library, or Skia bindings |
| A TLS library | Nearly every page is HTTPS | Python's ssl, rustls, or OpenSSL |
| Comfort with sockets | You'll speak HTTP over a raw TCP connection | Chapter 10 |
| The two key specs, bookmarked | Answers to "what should happen here?" | The WHATWG HTML Standard's parsing section; CSS 2.1 chapters 8 to 10 |
| A few saved test pages | Reproducible bugs | Save some simple pages locally, plus one big real one |
04The roadmap
Eight milestones. The first two give you something on screen quickly; the rest replace each crude stage with a real one.
Fetch a page by hand
1 eveningSplit the URL into scheme, host, port and path. Open a TCP socket, wrap it in
TLS for https, and write the request yourself: a request line, a Host
header, and a blank line. Read the status line and headers, then the body.
Getting the body right is the first real lesson. With keep-alive the server
won't close the connection, so you have to stop at Content-Length bytes or
decode chunked transfer encoding. Add redirects while you're here; a lot of
the web answers with a 301 first.
browse https://example.org prints the page's HTML, and it works for a server that sends a chunked body.Text in a window
1 eveningThrow the tags away for now: keep only the text between them, and decode the
common entities like & and <. Draw it word by word, measuring each
word with the font, and wrap when the next one won't fit.
Then add scrolling by moving everything up and skipping words that are off screen. It's crude, but it's a browser. Every milestone from here replaces one piece of this with the real thing.
HTML into a DOM
1–2 weekendsWrite a tokenizer that emits start tags with attributes, end tags, text and
comments. Then a tree builder that keeps a stack of open elements. Handle
void elements like br and img, which never get an end tag.
Real HTML is full of mistakes, and the spec says exactly how to repair them.
A p isn't closed before the next p, so the parser has to close it for you.
Implement the handful of implied end tags and implicit html, head and
body elements, and most pages will come out right.
CSS and the cascade
1–2 weekendsParse stylesheets into rules: a list of selectors and a list of declarations. Start with tag, class, ID and descendant selectors. Match them against each element, and when several set the same property, the most specific one wins, with source order breaking ties.
Then handle inheritance: properties like color and font-size pass down
to children unless something overrides them. You'll also want a small
user-agent stylesheet of your own, the defaults that make h1 big and p
a block with margins.
Block layout
1 weekendBuild a layout tree from the styled DOM. For a block box, the width comes from the parent's content width minus margins, borders and padding; the children are laid out top to bottom; and the height is the sum of theirs. That top-down, bottom-up pass is the heart of CSS layout.
Test it against a real browser. Make a page of coloured, nested boxes, open it in both, and compare. Matt Brubeck's "Let's build a browser engine" covers this milestone particularly clearly.
Inline layout and text
1–2 weekendsText isn't laid out as boxes stacked vertically. It flows into line boxes: each word is placed after the last, a new line starts when the next word won't fit, and every word on the line is aligned on a shared baseline using the font's ascent and descent.
Mixed content is where it gets tricky. When a block holds both text and block-level children, you need anonymous block boxes to wrap the loose text. Expect this milestone to take a bit longer than you planned.
Paint, scroll and follow links
1 weekendWalk the layout tree and emit a display list: rectangles for backgrounds and borders, text runs at positions. Drawing is now a loop over that list, and scrolling only changes an offset, so you don't re-run layout on every key press.
For clicks, do hit testing: find the deepest layout box under the pointer,
walk up to its a element, and resolve its href against the page URL.
Relative URLs, ../ segments and scheme-relative links all have rules. Keep a
history stack and you have a browser you could actually use for simple sites.
Stretch: run JavaScript
2–3 weekendsDon't write the language. Embed an existing engine such as QuickJS or
Duktape, run each script tag, and expose a small DOM API to it:
querySelector, addEventListener, innerHTML and textContent.
The interesting part is the other direction. A script can change the DOM at any point, so your browser has to mark what's stale and redo style, layout and paint before the next frame. This is where keeping the stages separate pays for itself.
querySelector and sets innerHTML on click changes the text, and your browser re-styles and re-lays out only after the handler returns.05Traps that catch everyone
| Symptom | Cause | Fix |
|---|---|---|
| The request hangs after the headers arrive | Waiting for the server to close a keep-alive connection | Read exactly Content-Length bytes, decode chunked bodies, or send Connection: close |
| The body is binary junk | You advertised Accept-Encoding: gzip and didn't decompress | Don't send that header, or decompress |
| Accented characters are garbled | Wrong character encoding | Honour the charset in Content-Type or the meta tag; default to UTF-8 |
| Every paragraph is nested inside the previous one | Unclosed p and li tags never get implied end tags | Close them when a sibling starts, as the spec's tree builder does |
| A rule seems to be ignored | Declarations sorted by order only | Sort by specificity first, then source order |
| All the text runs together on one line | Inline and block content treated the same | Separate block and inline formatting, with anonymous boxes |
| Scrolling is slow on long pages | Layout re-runs on every scroll | Scroll by offsetting the display list; skip items outside the viewport |
06Stretch goals
- Images. Decode PNGs or use a library, and give replaced elements an intrinsic size in layout.
- Forms and cookies. Submit a form with
POST, store cookies per origin, and log in to a simple site. You'll meet the same-origin policy quickly. - An HTTP cache. Honour
Cache-ControlandETag, and see how much faster a second visit gets. Chapter 25 covers the ideas. - Flexbox. A second layout algorithm shows you how layout modes plug into one engine.
- A compositor thread. Move raster and scrolling off the main thread so scrolling stays smooth while JavaScript runs.
07References worth your time
The free online book at browser.engineering. It builds a real browser in Python, chapter by chapter, and follows roughly the same road as this guide.
A short blog series building a toy engine called robinson in Rust. Its parsing, style and block layout posts are clear and quick to read.
The exact tokenizer states and tree-construction rules every browser follows. Read the parts you need, when you need them.
The box model and visual formatting model: the rules behind milestones 5 and 6.
A long article on web.dev tracing a page from network to pixels in real engines. A good map before you start.
An illustrated four-part series on Chrome's process model, rendering pipeline and compositor. Useful for the stretch goals.
An independent open-source browser engine written from scratch in C++, following the specs closely. Read its HTML tokenizer after milestone 3.