KnowSys
⇅Build it yourself

Build a TCP/IP stack in userspace

A network stack that runs as an ordinary process on a TAP device: Ethernet, ARP, IPv4, ICMP and UDP, then a TCP that handshakes, retransmits, respects windows and backs off under congestion, until curl can fetch a page through it.

Serious side project⏱ 8–12 weekendsC · Rust · Go

A connect() call looks like one step. Underneath, the kernel resolves a MAC address, builds headers for three protocols, computes checksums, runs a handshake, and then spends the life of the connection guessing how fast it may send without drowning the network. Almost none of that is visible from the socket API, and all of it shows up in production as timeouts, resets and throughput you can't explain.

This project moves the stack out of the kernel and into a program you write. Linux gives you a virtual network card, a TAP device, that hands you raw Ethernet frames as bytes and accepts frames back. From there you build each layer: answer ARP, reply to ping, echo UDP, and finally run TCP well enough that a real client on the host completes a handshake with you and downloads a file.

01Why build this

TCP is the protocol most engineers depend on and fewest have read. Building it pays off in ordinary debugging:

  • Packet captures become readable. After writing the header parser, a Wireshark trace of SYN, SYN-ACK, dup ACKs and retransmits reads like a log of decisions you've coded yourself.
  • Connection states mean something. TIME_WAIT piling up on a load balancer, CLOSE_WAIT leaking in a service: you'll know which side is stuck and why.
  • Throughput limits get explicable. Window size, round-trip time and packet loss decide how fast a connection goes. You'll have implemented each one.
  • Retries and timeouts in your own systems improve. TCP's retransmission timer is a small, well-studied answer to "how long should I wait before trying again?"

It's the hardest project in this list. Still, each layer below TCP works on its own, and it's pretty satisfying to see real tools answered by code you wrote.

02What you're building

An incoming TCP segment's path through the finished stack:

What your stack does when a segment arrives
●
⇩
TAP
read frame
▭
Ethernet
demux
◇
IPv4
validate
⇄
TCP
state machine
▤
Buffers
send / receive
▶
App
socket API
Step 1. read() on the TAP file descriptor returns one Ethernet frame: destination MAC, source MAC, EtherType, payload.
1 / 6

?Why does TCP need so much state per connection?

Because IP promises almost nothing. Packets can be lost, duplicated, reordered or delayed, and the other end can vanish without a word. TCP builds a reliable, ordered byte stream on top by numbering every byte, remembering what's been sent and acknowledged, and keeping timers for everything it's waiting on. Each piece of state exists to cover one specific way IP can let you down.

03Before you start

You needWhyWhere to get it
A Linux machine or VM with rootTAP devices are a Linux kernel featureAny current distro; a VM is safest
ip, tcpdump or Wireshark, ping, nc, curlYour test clients and your debuggerDistro packages
RFC 826, 791, 792, 768 and 9293The exact byte layouts and rulesrfc-editor.org
Comfort with bit fields and byte orderEvery header is big-endian with packed fieldsChapter 10 for the kernel's side
A systems language with precise control of bytesYou'll serialise headers by handC matches the RFCs' examples; Rust catches more mistakes

04The roadmap

Nine milestones. Milestones 1 to 4 take roughly an evening or two each. TCP starts at milestone 5 and probably takes most of your time.

1

Frames off the wire

1 evening

Open /dev/net/tun, create a TAP interface with ioctl(TUNSETIFF), bring it up with an address on the host, and read from the file descriptor in a loop. Each read returns one Ethernet frame.

A TUN device would hand you IP packets and skip Ethernet. Use TAP anyway: ARP only exists at layer 2, and the host won't send you any IP until you answer it.

You’ll learnTUN vs TAP/dev/net/tunIFF_TAP | IFF_NO_PIEtherType
Done when: Your program prints the source MAC and EtherType of every frame while ping runs against its address.
2

Answer ARP

1 evening

When the host wants to reach your IP, it broadcasts "who has this address?". Parse the ARP packet, and if the target is you, reply with your MAC. Keep a small ARP cache of senders you've seen, because you'll need their MACs to send anything back.

Byte order bites here first. Every multi-byte field is big-endian on the wire. Wrap the conversions in one place, so you don't scatter byte swaps through the code.

You’ll learnARP request / replyMAC addressesARP cachenetwork byte order
Done when: ip neigh on the host lists your stack's IP with the MAC you chose.
3

IPv4 and ping

1 weekend

Parse the IPv4 header, check its checksum, and drop anything not addressed to you. For ICMP echo requests, swap source and destination, change the type to echo reply, recompute both checksums, and send it back inside a new Ethernet frame.

The Internet checksum from RFC 1071 is a one's-complement sum of 16-bit words. It's a few lines, and it's used again in UDP and TCP, so write it once and test it against a captured packet.

You’ll learnIPv4 headerInternet checksumICMP echoTTL
Done when: ping from the host gets replies from your stack, and Wireshark shows no checksum errors.
4

UDP

1 evening

UDP adds only ports, a length and a checksum. The checksum covers a pseudo-header made from the IP addresses as well as the UDP header, which is the first time a layer reaches down into the one below it.

Build a tiny API here: bind a port, receive a datagram, send one. You'll reuse the shape for TCP, and it keeps the protocol code apart from the application.

You’ll learnportspseudo-header checksumdemultiplexinga socket-like API
Done when: nc -u from the host sends a line to your stack and gets it echoed back.
5

The TCP handshake

1–2 weekends

Keep a transmission control block per connection, keyed by the four-tuple, holding the send and receive sequence variables from RFC 9293 (SND.UNA, SND.NXT, RCV.NXT and friends). On a SYN to a listening port, pick an initial sequence number and reply SYN-ACK. On the final ACK, move to ESTABLISHED.

Send RST for segments to ports nobody is listening on. Compare sequence numbers with wraparound arithmetic from the start: they're 32-bit and they wrap.

You’ll learnsequence numbersSYN / SYN-ACK / ACKthe TCBRST
Done when: nc on the host connects to your stack, and ss -t shows the connection as ESTAB.
6

Moving data

1–2 weekends

Accept in-order data into a receive buffer, advance RCV.NXT, and ACK it. Queue out-of-order segments until the gap fills, instead of dropping them. On the send side, split application data into segments no larger than the peer's MSS and within its advertised window.

A tiny HTTP responder on top gives you a real client to test against. Once curl works, capture the exchange and compare it with the same request to a kernel socket.

You’ll learnbyte-stream ACKsreceive windowout-of-order segmentsPSH
Done when: curl on the host fetches a page from a tiny HTTP responder running on your stack.
7

Closing, all eleven states

1 weekend

Implement FIN in both directions: the side that closes first goes through FIN_WAIT_1, FIN_WAIT_2 and TIME_WAIT, the other through CLOSE_WAIT and LAST_ACK. Each direction closes independently, so one side can keep sending after the other is done.

Draw the state diagram from RFC 9293 and put it next to your code. Most TCP bugs at this stage are a transition the diagram has and your match or switch doesn't.

You’ll learnFIN handshakehalf-closeTIME_WAITsimultaneous close
Done when: Connections closed from either side reach CLOSED on both, and your stack holds TIME_WAIT before reusing a four-tuple.
8

Retransmission

1–2 weekends

Keep unacknowledged segments and a retransmission timer. Estimate round-trip time the way RFC 6298 describes, with a smoothed average and a variance term, and set the timeout from both. Don't sample RTT from retransmitted segments, since you can't tell which copy the ACK is for.

Double the timeout on every consecutive retransmission. Then use tc qdisc ... netem to add loss, delay and reordering, and watch your stack recover.

You’ll learnRTOSRTT and RTTVARKarn's algorithmexponential backoff
Done when: With tc netem dropping 10% of packets on the TAP interface, a file transfer still completes byte for byte.
9

Flow and congestion control

2 weekends

Respect the receiver's window, including a zero window: stop sending and probe periodically until it opens. Then add a congestion window from RFC 5681. Start small and grow it quickly in slow start, grow linearly after a threshold, cut it on loss.

On three duplicate ACKs, retransmit right away instead of waiting for the timer. Log the congestion window on every ACK and plot it. That sawtooth is a better picture of how TCP behaves than any diagram.

You’ll learnslow startcongestion avoidancefast retransmitzero-window probes
Done when: Under induced loss, a plot of your congestion window shows slow start, a sawtooth, and a halving after three duplicate ACKs.

05Traps that catch everyone

SymptomCauseFix
The host never sends you IP packetsARP isn't answered, or your IP equals the TAP interface'sAnswer ARP first; use a separate address on the same subnet
Wireshark flags every packet's checksumWrong byte order, odd-length payload not padded, or the checksum field not zeroed firstZero the field, pad to 16 bits, test against a captured packet
The host sends RST to your SYN-ACKYou're using raw sockets and the kernel doesn't know the connectionUse a TAP device so the kernel stack stays out of the way
Long transfers break at a seemingly random pointSequence numbers compared with plain <, which fails once they wrapCompare using signed 32-bit differences
Transfers stall under lossNo retransmission of the last segment, or timer never restartedRestart the timer whenever new data is ACKed and something is still in flight
Throughput collapses on reorderingEach out-of-order segment was droppedQueue out-of-order data and send an ACK that tells the peer what's missing
Stuck in CLOSE_WAITYour app never closed its sideCLOSE_WAIT means the local application hasn't called close; the fix is in the app

06Stretch goals

  • A real socket API. Expose your stack through a library with connect, accept, read and write, and run an unmodified program on it.
  • TCP options. MSS, window scaling and timestamps from RFC 7323, then SACK from RFC 2018.
  • A modern congestion controller. Implement CUBIC and compare its window graph with Reno's under the same loss.
  • IPv6 and NDP. Replace ARP with Neighbor Discovery and see how much of the stack carries over.
  • DHCP and DNS clients, so your stack can join a real network and resolve names without help from the host.

07References worth your time

saminiir, Let's code a TCP/IP stack

A five-part blog series building a userspace stack in C on a TAP device, from Ethernet and ARP to TCP retransmission. The closest match to this roadmap.

RFC 9293, Transmission Control Protocol

The 2022 consolidation of RFC 793 and its many updates. Keep the state diagram and the segment-arrival rules open while you code.

Kevin Fall and W. Richard Stevens, TCP/IP Illustrated, Volume 1

The protocol suite explained through real packet traces. The TCP chapters are the best companion to milestones 5 to 9.

Jon Gjengset, Implementing TCP in Rust

A long livestream series writing a TCP implementation on a TUN device, reading RFC 793 line by line. Good for seeing the debugging, not just the result.

RFC 6298 and RFC 5681

Computing the retransmission timer, and TCP congestion control. Both are short and precise.

smoltcp and lwIP

Two small, production-quality TCP/IP stacks for embedded systems, in Rust and C. Read them to see how real code organises the layers you built.

Chapters that back this project

Next project⬡ a replicated key-value store with Raft→