A connect() call looks like one step. Underneath, the kernel resolves a MAC
address, builds headers for three protocols, computes checksums, runs a
handshake, and then spends the life of the connection guessing how fast it may
send without drowning the network. Almost none of that is visible from the
socket API, and all of it shows up in production as timeouts, resets and
throughput you can't explain.
This project moves the stack out of the kernel and into a program you write.
Linux gives you a virtual network card, a TAP device, that hands you raw
Ethernet frames as bytes and accepts frames back. From there you build each
layer: answer ARP, reply to ping, echo UDP, and finally run TCP well enough
that a real client on the host completes a handshake with you and downloads a
file.
01Why build this
TCP is the protocol most engineers depend on and fewest have read. Building it pays off in ordinary debugging:
- Packet captures become readable. After writing the header parser, a Wireshark trace of SYN, SYN-ACK, dup ACKs and retransmits reads like a log of decisions you've coded yourself.
- Connection states mean something.
TIME_WAITpiling up on a load balancer,CLOSE_WAITleaking in a service: you'll know which side is stuck and why. - Throughput limits get explicable. Window size, round-trip time and packet loss decide how fast a connection goes. You'll have implemented each one.
- Retries and timeouts in your own systems improve. TCP's retransmission timer is a small, well-studied answer to "how long should I wait before trying again?"
It's the hardest project in this list. Still, each layer below TCP works on its own, and it's pretty satisfying to see real tools answered by code you wrote.
02What you're building
An incoming TCP segment's path through the finished stack:
read() on the TAP file descriptor returns one Ethernet frame: destination MAC, source MAC, EtherType, payload.?Why does TCP need so much state per connection?
Because IP promises almost nothing. Packets can be lost, duplicated, reordered or delayed, and the other end can vanish without a word. TCP builds a reliable, ordered byte stream on top by numbering every byte, remembering what's been sent and acknowledged, and keeping timers for everything it's waiting on. Each piece of state exists to cover one specific way IP can let you down.
03Before you start
| You need | Why | Where to get it |
|---|---|---|
| A Linux machine or VM with root | TAP devices are a Linux kernel feature | Any current distro; a VM is safest |
ip, tcpdump or Wireshark, ping, nc, curl | Your test clients and your debugger | Distro packages |
| RFC 826, 791, 792, 768 and 9293 | The exact byte layouts and rules | rfc-editor.org |
| Comfort with bit fields and byte order | Every header is big-endian with packed fields | Chapter 10 for the kernel's side |
| A systems language with precise control of bytes | You'll serialise headers by hand | C matches the RFCs' examples; Rust catches more mistakes |
04The roadmap
Nine milestones. Milestones 1 to 4 take roughly an evening or two each. TCP starts at milestone 5 and probably takes most of your time.
Frames off the wire
1 eveningOpen /dev/net/tun, create a TAP interface with ioctl(TUNSETIFF), bring it
up with an address on the host, and read from the file descriptor in a loop.
Each read returns one Ethernet frame.
A TUN device would hand you IP packets and skip Ethernet. Use TAP anyway: ARP only exists at layer 2, and the host won't send you any IP until you answer it.
ping runs against its address.Answer ARP
1 eveningWhen the host wants to reach your IP, it broadcasts "who has this address?". Parse the ARP packet, and if the target is you, reply with your MAC. Keep a small ARP cache of senders you've seen, because you'll need their MACs to send anything back.
Byte order bites here first. Every multi-byte field is big-endian on the wire. Wrap the conversions in one place, so you don't scatter byte swaps through the code.
ip neigh on the host lists your stack's IP with the MAC you chose.IPv4 and ping
1 weekendParse the IPv4 header, check its checksum, and drop anything not addressed to you. For ICMP echo requests, swap source and destination, change the type to echo reply, recompute both checksums, and send it back inside a new Ethernet frame.
The Internet checksum from RFC 1071 is a one's-complement sum of 16-bit words. It's a few lines, and it's used again in UDP and TCP, so write it once and test it against a captured packet.
ping from the host gets replies from your stack, and Wireshark shows no checksum errors.UDP
1 eveningUDP adds only ports, a length and a checksum. The checksum covers a pseudo-header made from the IP addresses as well as the UDP header, which is the first time a layer reaches down into the one below it.
Build a tiny API here: bind a port, receive a datagram, send one. You'll reuse the shape for TCP, and it keeps the protocol code apart from the application.
nc -u from the host sends a line to your stack and gets it echoed back.The TCP handshake
1–2 weekendsKeep a transmission control block per connection, keyed by the four-tuple,
holding the send and receive sequence variables from RFC 9293 (SND.UNA,
SND.NXT, RCV.NXT and friends). On a SYN to a listening port, pick an initial
sequence number and reply SYN-ACK. On the final ACK, move to ESTABLISHED.
Send RST for segments to ports nobody is listening on. Compare sequence numbers with wraparound arithmetic from the start: they're 32-bit and they wrap.
nc on the host connects to your stack, and ss -t shows the connection as ESTAB.Moving data
1–2 weekendsAccept in-order data into a receive buffer, advance RCV.NXT, and ACK it.
Queue out-of-order segments until the gap fills, instead of dropping them. On
the send side, split application data into segments no larger than the peer's
MSS and within its advertised window.
A tiny HTTP responder on top gives you a real client to test against. Once
curl works, capture the exchange and compare it with the same request to a
kernel socket.
curl on the host fetches a page from a tiny HTTP responder running on your stack.Closing, all eleven states
1 weekendImplement FIN in both directions: the side that closes first goes through FIN_WAIT_1, FIN_WAIT_2 and TIME_WAIT, the other through CLOSE_WAIT and LAST_ACK. Each direction closes independently, so one side can keep sending after the other is done.
Draw the state diagram from RFC 9293 and put it next to your code. Most TCP
bugs at this stage are a transition the diagram has and your match or
switch doesn't.
Retransmission
1–2 weekendsKeep unacknowledged segments and a retransmission timer. Estimate round-trip time the way RFC 6298 describes, with a smoothed average and a variance term, and set the timeout from both. Don't sample RTT from retransmitted segments, since you can't tell which copy the ACK is for.
Double the timeout on every consecutive retransmission. Then use tc qdisc ... netem to add loss, delay and reordering, and watch your stack recover.
tc netem dropping 10% of packets on the TAP interface, a file transfer still completes byte for byte.Flow and congestion control
2 weekendsRespect the receiver's window, including a zero window: stop sending and probe periodically until it opens. Then add a congestion window from RFC 5681. Start small and grow it quickly in slow start, grow linearly after a threshold, cut it on loss.
On three duplicate ACKs, retransmit right away instead of waiting for the timer. Log the congestion window on every ACK and plot it. That sawtooth is a better picture of how TCP behaves than any diagram.
05Traps that catch everyone
| Symptom | Cause | Fix |
|---|---|---|
| The host never sends you IP packets | ARP isn't answered, or your IP equals the TAP interface's | Answer ARP first; use a separate address on the same subnet |
| Wireshark flags every packet's checksum | Wrong byte order, odd-length payload not padded, or the checksum field not zeroed first | Zero the field, pad to 16 bits, test against a captured packet |
| The host sends RST to your SYN-ACK | You're using raw sockets and the kernel doesn't know the connection | Use a TAP device so the kernel stack stays out of the way |
| Long transfers break at a seemingly random point | Sequence numbers compared with plain <, which fails once they wrap | Compare using signed 32-bit differences |
| Transfers stall under loss | No retransmission of the last segment, or timer never restarted | Restart the timer whenever new data is ACKed and something is still in flight |
| Throughput collapses on reordering | Each out-of-order segment was dropped | Queue out-of-order data and send an ACK that tells the peer what's missing |
| Stuck in CLOSE_WAIT | Your app never closed its side | CLOSE_WAIT means the local application hasn't called close; the fix is in the app |
06Stretch goals
- A real socket API. Expose your stack through a library with
connect,accept,readandwrite, and run an unmodified program on it. - TCP options. MSS, window scaling and timestamps from RFC 7323, then SACK from RFC 2018.
- A modern congestion controller. Implement CUBIC and compare its window graph with Reno's under the same loss.
- IPv6 and NDP. Replace ARP with Neighbor Discovery and see how much of the stack carries over.
- DHCP and DNS clients, so your stack can join a real network and resolve names without help from the host.
07References worth your time
A five-part blog series building a userspace stack in C on a TAP device, from Ethernet and ARP to TCP retransmission. The closest match to this roadmap.
The 2022 consolidation of RFC 793 and its many updates. Keep the state diagram and the segment-arrival rules open while you code.
The protocol suite explained through real packet traces. The TCP chapters are the best companion to milestones 5 to 9.
A long livestream series writing a TCP implementation on a TUN device, reading RFC 793 line by line. Good for seeing the debugging, not just the result.
Computing the retransmission timer, and TCP congestion control. Both are short and precise.
Two small, production-quality TCP/IP stacks for embedded systems, in Rust and C. Read them to see how real code organises the layers you built.