Chapter 12 · Part IV · Routing

VPNs, tunnels & overlays

A tunnel is a packet carried inside another packet. That one trick gives you corporate VPNs, encrypted links across hostile Wi-Fi, and mesh networks like Tailscale where every device can reach every other as if they shared a switch. It also costs bytes, shrinks the MTU, adds virtual adapters, and rewrites the routing table. This chapter is about both halves.

60 bytesWireGuard's overhead per packet over IPv4: 20 IP + 8 UDP + 16 header + 16 auth tag. Over IPv6 it's 80
100.64.0.0/10the shared address space (RFC 6598) Tailscale draws device addresses from: 4,194,304 addresses
1280bytes: the IPv6 minimum MTU, and the MTU Tailscale gives its interface by default so it fits through nearly any path

Common mix-up: when Tailscale says a connection is going through a relay, your traffic is not exposed. The DERP relay forwards packets that are already WireGuard-encrypted end to end, and the coordination server never carries traffic at all. A relay costs latency and throughput, not privacy.

The trick

A packet inside a packet

Every network in this book moves packets by looking at their headers. A router reads the destination IP, looks it up, and sends the packet on. It never looks at what's inside. That indifference is what makes tunnels possible: take a complete packet, headers and all, treat it as ordinary data, and put a new header in front of it. To every router along the way, the result is just a packet addressed from one tunnel endpoint to the other. The original packet rides along as cargo.

That act of wrapping is encapsulation. At the far end, the receiving endpoint strips the outer header, finds a perfectly ordinary packet inside, and hands it to its own routing table as if it had arrived on a local cable. The two networks at either end now behave as though they were directly connected, even though there might be fifteen routers and an ocean between them. Networking people call the network you're simulating the overlay, and the real one carrying it the underlay.

The simplest tunnels don't even encrypt. IP-in-IP (RFC 2003) adds a single 20-byte IPv4 header. GRE (RFC 2784) adds that plus a 4-byte header naming what's inside, so it can carry IPv6 or other protocols too. VXLAN (RFC 7348), the workhorse of data centers, wraps a whole Ethernet frame in UDP so that virtual machines in different buildings can share one Layer 2 segment. None of these hide anything; anyone on the path can read the inner packet.

A VPN, a virtual private network, is a tunnel that also encrypts and authenticates. The inner packet becomes unreadable ciphertext, and a cryptographic tag proves it wasn't changed on the way. Now the café's Wi-Fi, the hotel's router and your ISP see only a stream of UDP packets between your laptop and one far-off address. They can tell how much you're sending and when, but not to whom or what.

on the wire = outer headers + inner packet + trailers
Every tunnel in this chapter is this sum with different numbers in it. The inner packet is untouched; it's the envelope that changes.
TunnelWhat it adds (outer IPv4)BytesEncrypted?
IP-in-IPIPv4 header20no
GREIPv4 + GRE header24no
VXLANIPv4 + UDP + VXLAN + inner Ethernet header50no
WireGuardIPv4 + UDP + WireGuard header + Poly1305 tag60yes
IPsec ESP, tunnel mode, AES-GCMIPv4 + ESP header + IV + padding + trailer + ICV~54–57yes
…plus NAT traversalthe same inside an extra UDP header (port 4500)~62–65yes
Go deeper: TUN, TAP, and which layer you're tunneling

A tunnel can carry Layer 3 (IP packets) or Layer 2 (whole Ethernet frames). The operating system needs a fake network adapter for each. A TUN adapter (for "tunnel") exchanges raw IP packets with the VPN program: the kernel routes a packet to it, and instead of going to a wire, the packet is handed to user space to be encrypted and sent. A TAP adapter (for "network tap") does the same with Ethernet frames, MAC addresses and all, which lets broadcast traffic and non-IP protocols cross, at the price of 14 more bytes per packet and a lot more chatter.

WireGuard and Tailscale use TUN. On Windows that's Wintun, a small TUN driver written by the WireGuard project. OpenVPN historically used TAP-Windows, which can run in either mode. VXLAN is the Layer 2 case on servers: the inner Ethernet header is why its overhead is 50 rather than 36.

The cost

The MTU squeeze

Every link has a maximum transmission unit, the biggest IP packet it will carry in one frame. On Ethernet and Wi-Fi it's 1,500 bytes. On a DSL line using PPPoE it's 1,492, because PPPoE takes 8 bytes of every frame for itself. Applications don't usually think about this; TCP, which carries most traffic, quietly sizes its segments to fill 1,500-byte packets.

Now put a tunnel on that 1,500-byte link. A full 1,500-byte packet goes into WireGuard, gains 60 bytes of envelope, and comes out at 1,560. That doesn't fit. Something has to give, and there are only three things that can happen:

  • Fragmentation. If the outer packet is IPv4 without the Don't Fragment bit, a router splits it into two pieces, here a 1,500-byte fragment and an 80-byte one, and the far end glues them back together. It works, but every full-size packet becomes two, the receiver has to buffer and reassemble, and losing either piece loses both.
  • Drop and tell. If Don't Fragment is set (or the outer packet is IPv6, where routers never fragment), the router drops it and sends back an ICMP message: "Fragmentation Needed" in IPv4, "Packet Too Big" in IPv6, carrying the MTU that would have worked. The sender shrinks its packets and tries again. This is path MTU discovery (RFC 1191 for IPv4, RFC 8201 for IPv6).
  • Drop and stay silent. If a firewall somewhere blocks those ICMP messages, which is depressingly common, the sender never learns. Small packets go through, big ones vanish. This is the classic PMTU black hole: a web page starts loading and stalls, SSH logs in fine and then hangs the moment you run a command with a lot of output, a file copy starts and freezes at zero.

The clean fix is to make the tunnel's virtual adapter honest about its size. If the underlay carries 1,500 and the tunnel adds 60, give the tunnel interface an MTU of 1,440. Then the operating system never hands the tunnel a packet that won't fit after wrapping. wg-quick, the standard WireGuard setup tool, defaults to 1,420: it subtracts 80 to cover the worse case of an IPv6 outer header. Tailscale goes further and uses 1,280, the minimum every IPv6 link must support, trading a few percent of efficiency for a size that fits through almost anything, including tunnels inside tunnels.

TCP gets one more safeguard. When a connection opens, each side announces a maximum segment size, the most data it wants in one segment: the MTU minus 40 bytes of IPv4 and TCP headers (60 for IPv6). A host behind a 1,440-byte tunnel announces 1,400. Routers and VPN gateways can also rewrite this value in passing SYN packets, called MSS clamping, which rescues TCP even when the endpoints don't know a tunnel is there.

tunnel MTU = link MTU − overhead  ·  MSS = tunnel MTU − 40
For WireGuard over IPv4 on Ethernet: 1500 − 60 = 1440, and a TCP MSS of 1400. For IPv6 inside, subtract 60 instead of 40.
Go deeper: why TCP inside TCP is a bad idea, and what PLPMTUD fixes

Some VPNs can run over TCP, OpenVPN's TCP mode on port 443 being the usual example, because TCP on 443 slips through firewalls that block everything else. It works until there's packet loss. Then the outer TCP connection stalls and retransmits, while every inner TCP connection, which can't see the outer one's recovery, also times out and retransmits, and the two layers of backoff stack on each other. Throughput collapses far more than the loss rate would suggest. That's why WireGuard is UDP only and OpenVPN prefers UDP.

Because ICMP black holes are so common, newer stacks can find the path MTU without ICMP at all. Packetization Layer PMTUD (RFC 4821, and RFC 8899 for datagram protocols) sends probes of increasing size within the connection itself and watches which ones get acknowledged. Linux TCP will fall back to it when it suspects a black hole (the tcp_mtu_probing setting). RFC 4459 is the standards body's own catalog of the MTU headaches tunnels cause.

Instrument 1

Encapsulation & MTU calculator

Stack tunnels on a link and send a packet through them. The bar is the packet as it crosses the wire, outermost header on the left; the dashed line is the link's MTU. The calculator wraps the packet one layer at a time, working out IPsec's alignment padding as it goes, then finds the largest inner packet that still fits, the TCP MSS that goes with it, and what happens to a packet that's too big.

Total overhead–
Safe inner MTU–
TCP MSS (IPv4 · IPv6)–
This packet–

Three families

IPsec, OpenVPN, WireGuard

Almost every VPN you'll meet is one of three designs. They all encrypt IP packets and send them to a peer; they differ in how the two ends agree on keys, what they ride on, and how much machinery comes with them.

IPsec is the standards-track veteran, built into Windows, macOS, iOS, Android, Linux and nearly every firewall appliance. It splits the work in two. IKEv2 (RFC 7296) is the negotiation: the peers authenticate with certificates or a shared secret, agree on algorithms, and derive keys, over UDP port 500. ESP (RFC 4303) then carries the encrypted packets as IP protocol 50, a protocol of its own rather than TCP or UDP. That last detail is a problem behind NAT, which needs port numbers to tell flows apart, so when IKE detects a NAT in the path, both sides move to UDP port 4500 and wrap ESP inside UDP (RFC 3948, NAT traversal). IPsec is powerful and interoperable, and its configuration space is enormous: dozens of algorithm combinations, two modes, two protocols, policies that decide which traffic gets protected.

OpenVPN is an open-source program, not a standard. Its control channel is TLS, the same protocol that secures web pages, which makes certificates and usernames easy. Its data channel runs over UDP (port 1194 by default) or TCP, through a TUN or TAP adapter, entirely in user space. It's flexible, works through nearly any firewall, and has been the default choice for commercial VPN services and self-hosted remote access for twenty years. It's also comparatively slow, because every packet makes a round trip between the kernel and a user-space process.

WireGuard is the newcomer, merged into the Linux kernel in 2020. Its design is deliberately small: the Linux implementation is around 4,000 lines of code, against tens or hundreds of thousands for typical IPsec and OpenVPN implementations. There's no negotiation of algorithms, just one fixed modern set: Curve25519 for key exchange, ChaCha20-Poly1305 for encryption, BLAKE2s for hashing, wrapped in a handshake built on the Noise framework. Each peer is identified by a public key, the way an SSH server is. It runs over UDP only, and it's silent: a WireGuard endpoint doesn't answer packets that don't authenticate, so to a port scanner it looks like nothing is there.

Two WireGuard ideas matter for the rest of the chapter. The first is cryptokey routing: each peer's configuration lists AllowedIPs, the inner addresses that peer is allowed to send from and that should be sent to it. That one list is both the routing table and the access control list of the tunnel. The second is roaming: a peer's outer address isn't fixed. Whenever a correctly authenticated packet arrives from a new address and port, WireGuard simply updates where it sends replies. Walk from Wi-Fi to cellular and the tunnel follows you without renegotiating anything.

IPsec (IKEv2 + ESP)OpenVPNWireGuard
Identitycertificates, EAP, or pre-shared keyTLS certificates, often plus username/passwordone Curve25519 public key per peer
TransportUDP 500 for IKE; ESP is IP protocol 50, or UDP 4500 behind NATUDP 1194 by default, or TCPUDP only, any port (51820 is customary)
Algorithmsnegotiated from a long listnegotiatedfixed; a new version would change them all at once
Where it runskerneluser spacekernel on Linux; fast user-space ports elsewhere
RoamingMOBIKE extension (RFC 4555)reconnect, or float optionbuilt in: replies follow the latest authenticated source
Adapternone needed on many platforms (policy-based)TUN or TAPTUN (Wintun on Windows)
Go deeper: how WireGuard stays quiet and still gets through NAT

WireGuard's handshake is a single round trip: an initiation message and a response, 148 and 92 bytes. Keys are refreshed by a new handshake every two minutes of use, so a stolen session key is good for very little traffic. In between, a peer that has nothing to send sends nothing. That's elegant, but it means a NAT's mapping for an idle tunnel will time out and the far side will lose the way back in (Chapter 8). The PersistentKeepalive option fixes it: an empty authenticated packet every N seconds, 25 being the usual choice, keeps the mapping alive.

Each data packet carries a 4-byte message type, a 4-byte receiver index (which session this is), an 8-byte counter used both as the nonce and to reject replays, then the ciphertext and a 16-byte Poly1305 tag. The plaintext is padded to a multiple of 16 bytes, but never past the interface MTU, so padding doesn't change the safe MTU in the calculator above.

Which packets go in

Full tunnel, split tunnel

A VPN connection is two things: an encrypted pipe, and a set of routes that decide what goes into it. The pipe is the same in every case. The routes are a policy choice, and they're where most of the confusion lives.

A full tunnel sends everything through the VPN. The VPN client adds routes that cover the whole internet, either a new default route with a better metric or, more robustly, the two half-internet routes 0.0.0.0/1 and 128.0.0.0/1, which beat any /0 by longest-prefix match (Chapter 10). Your web browsing, your DNS lookups and your game traffic all exit from the VPN server. That's what consumer privacy VPNs do and what a cautious company might require on public Wi-Fi.

A split tunnel sends only some prefixes through the VPN. The corporate VPN adds 10.0.0.0/8 and the company's public ranges; everything else uses your normal default route and goes straight to the internet. That's faster, it doesn't drag your video calls through the office firewall, and it keeps the VPN server from becoming a bottleneck. It's also exactly how Tailscale behaves out of the box: it adds a route for 100.64.0.0/10 (plus any subnet routes you've approved), and the rest of your traffic is untouched until you choose an exit node, which turns it into a full tunnel through one of your own machines.

Three details make or break both modes:

  • The pin route. The encrypted packets themselves are addressed to the VPN server. If the tunnel's routes cover that address, which a full tunnel always does, the tunnel would try to send its own transport into itself. So the client also adds a /32 for the server via the physical adapter, the most specific route in the table. Remove it in the instrument below and watch the tunnel swallow itself. (Clients on Linux often use policy routing with a firewall mark instead, which does the same job.)
  • The local network. Your LAN's connected route, 192.168.86.0/24, is a /24, longer than either /1, so even a full tunnel normally leaves the printer reachable. VPNs that want to forbid that ("block LAN access") have to do it with firewall rules, not routes.
  • DNS. Routing the packets isn't the same as routing the questions. If the VPN pushes its own DNS servers but your machine still asks the café's resolver first, every name you look up leaks, and internal names like intranet.corp fail to resolve. Split DNS, sending only certain domains to the VPN's resolver, is the split tunnel's twin.
Go deeper: private addresses that collide

Split tunnels assume the VPN's prefixes don't overlap your local ones. When they do, longest-prefix match picks a winner and the loser silently becomes unreachable. If the office uses 192.168.1.0/24 and so does your home router, a VPN route for 192.168.1.0/24 makes your own router vanish; a longer, more specific route on either side wins outright. This is one reason Tailscale picked addresses from 100.64.0.0/10: that block is reserved for carrier-grade NAT (RFC 6598), so it almost never collides with a home or office LAN. The rare exception is a laptop behind an ISP that really does use CGNAT addresses on its customer side.

Instrument 2

Split-tunnel route visualizer

Choose which prefixes go into the tunnel, pick a destination, and send a packet. The laptop runs a real longest-prefix-match lookup on the table below, then, if the packet goes into the tunnel, looks up the outer packet too, because the wrapped packet needs a route of its own. Teal packets are plain; violet packets are wrapped and encrypted.

Route chosen–
Path–
Your ISP sees–
Outcome–

      

Mesh

Overlays: every device a peer

A classic VPN is a hub: every laptop dials one concentrator, and two laptops that want to talk to each other go through it, even if they're sitting at the same desk. An overlay mesh flips that. Every device gets a WireGuard tunnel directly to every other device it's allowed to reach, and the "network" is the set of all those pairwise tunnels. Tailscale is the best-known example; ZeroTier, Nebula and Netmaker are relatives.

The hard parts of a mesh aren't the encryption, which WireGuard already does, but the bookkeeping. With N devices there are N(N−1)/2 possible pairs; each one needs the other's public key, current address and port, and permission. Tailscale splits that into a control plane and a data plane:

  • The coordination server is the control plane. Each device generates its own key pair, keeps the private key, and uploads only the public key, along with the addresses where it might be reached. The server checks who you are (through your identity provider), applies the access rules, and sends each device a map of the peers it may talk to. It never sees or carries traffic.
  • WireGuard tunnels are the data plane: device to device, encrypted with keys the server never had. Each device gets a stable address from 100.64.0.0/10, which stays the same as the device moves between networks, plus an IPv6 address from fd7a:115c:a1e0::/48.
  • NAT traversal is what makes "device to device" possible when both are behind home routers. Each device learns its public address and port by asking STUN servers (UDP 3478, RFC 5389), shares its candidates through the coordination server, and then both sides send packets at each other at once so that each NAT sees an outgoing flow and lets the replies in: hole punching (Chapter 9). It also asks the router directly for a port mapping when the router supports it, via UPnP, NAT-PMP (RFC 6886) or PCP (RFC 6887). Tailscale's own WireGuard port is UDP 41641 by default.
  • DERP relays are the fallback. DERP, Designated Encrypted Relay for Packets, is a network of Tailscale servers that forward already-encrypted WireGuard packets over HTTPS (TCP 443), which gets through almost any firewall. Every connection actually starts on DERP, so the first packets flow immediately, and moves to a direct path once hole punching succeeds. If it never succeeds, typically because a NAT assigns a different public port for every destination, traffic stays relayed.
  • MagicDNS gives every device a name. Tailscale runs a resolver at 100.100.100.100 on each machine that answers for names in your tailnet, so xps13 resolves to its 100.x address without anyone maintaining a DNS zone.

You can watch the data plane choose. tailscale status lists peers and marks each connection direct or relay; tailscale ping xps13 shows the first replies coming "via DERP(nyc)" and then, if hole punching works, switching to "via 192.168.86.40:41641". tailscale netcheck reports what the device has learned about its own NAT: whether UDP works, whether the mapping varies by destination, whether port mapping is available, and the latency to each DERP region.

direct if (hole punch succeeds) else DERP relay
Same encryption either way. The difference is the route: one hop across your LAN, or out to a relay and back.
Go deeper: why two devices on the same LAN can still end up relayed

Two devices on the same LAN should find each other trivially: each advertises its local address (192.168.86.x) as a candidate, and a direct packet across the switch needs no NAT at all. When that fails, it's usually one of three things. A firewall on one machine drops inbound UDP on the tunnel port. One machine advertises the wrong local candidates, for example an address on a Hyper-V or bridge adapter instead of the real LAN one. Or the client keeps throwing its path away: every time the operating system reports a network change, the client re-discovers its addresses, re-STUNs, and re-validates its paths, and traffic falls back to DERP while it does. A single change costs a second or two. A steady stream of them, from adapters that never stop coming and going, can keep a connection relayed indefinitely.

A double NAT, a home router behind another router (192.168.86.x behind 192.168.1.x), makes the internet-facing part harder: the outer NAT may not support port mapping, and NAT-PMP requests to the inner router get mapped only on the inner router. That matters for reaching the device from outside. For two devices on the same inner LAN, the direct path doesn't touch either NAT, so if they're relayed, look at the devices first.

The XPS case

Virtual adapters and link-change storms

Every tunnel, virtual machine switch and bridge on a computer shows up as a network adapter. Most are harmless one at a time. Collect enough and the computer's view of its own network never sits still, and the programs that care about that, overlay clients most of all, spend their time reacting.

  • VPN TUN/TAP adapters stay installed after you stop using the VPN. A disconnected TAP adapter sits at "media disconnected", but software that pokes at it, or a driver that wakes it on resume, can bring it up briefly, give it a self-assigned 169.254.x.x address because there's no DHCP behind it, and take it down again.
  • A Hyper-V external virtual switch takes over a physical adapter. The real NIC stops having an IP address; it becomes a port on a software switch, and the host gets a new adapter called vEthernet (name) that holds the address instead. Every virtual machine attached gets its own virtual NIC on the same switch. The separate Default Switch is a NAT network with its own private subnet, which Windows may choose differently after a reboot.
  • A Windows Network Bridge joins two adapters at Layer 2 and gives the host one more adapter, again taking IP configuration away from its members. Bridges and external vSwitches on the same NIC fight; neither is designed to coexist with the other.

Each of those events, an adapter appearing, disappearing, getting or losing an address, is broadcast by the operating system as a network change. The routing table gets rewritten (Chapter 10's adapter simulator counts the rewrites). And Tailscale's client, which watches for exactly those notifications with its link monitor, has to assume its paths may have changed: it re-reads interfaces, re-gathers candidate addresses, re-queries STUN, may retry port mapping on the router, and re-establishes paths to peers. If a change arrives every few seconds, that work never finishes, the client burns CPU doing it, and peers that were direct fall back to DERP in the meantime.

That's what was happening on J's XPS 13: twelve adapters, among them a Hyper-V external switch, a Network Bridge, and dead VPN TAP adapters, behind a double NAT with NAT-PMP requests that kept retrying. Remote Desktop to the machine over Tailscale was riding a DERP relay instead of the LAN, so every screen update made a trip to a relay server and back. The network wasn't broken. It was busy.

Hygiene stepWhyHow (Windows)
Remove dead TAP/TUN adaptersno more phantom up/down events or stray link-local addressesuninstall the old VPN; Device Manager → View → Show hidden devices → Network adapters; or pnputil /enum-devices /class Net then pnputil /remove-device
One external vSwitch, at mostone physical NIC, one owner, one host addressGet-VMSwitch; prefer an internal or Default Switch unless VMs truly need the LAN
No Network Bridge on topa bridge and a vSwitch both claim the NICNetwork Connections → right-click Network Bridge → Delete
Check the overlay's viewconfirm candidates and the path chosentailscale netcheck, tailscale status, tailscale ping <peer>
List what's leftsee hidden and disconnected adapters tooGet-NetAdapter -IncludeHidden
Go deeper: what "a network change" means to the operating system

Windows exposes changes through notification APIs such as NotifyIpInterfaceChange, NotifyUnicastIpAddressChange and NotifyRouteChange2; Linux sends netlink messages; macOS uses the System Configuration framework and routing sockets. A program subscribes and gets a callback per event. An adapter coming up can produce several in a row: the interface appears, gets a link-local IPv6 address, gets an IPv4 address, gains connected and /32 routes. Well-written clients debounce, waiting a moment and then evaluating the net effect, and try to tell a "major" change (the default route moved, the primary address changed) from noise. But a client can't know for certain that a change didn't matter until it looks, and looking is the expensive part.

Cheat sheet

Terms from this chapter

Encapsulation
Wrapping a complete packet as the payload of another packet.
Underlay / overlay
The real network that carries the tunnel / the virtual network the tunnel creates.
MTU
Largest IP packet a link carries in one frame. 1,500 on Ethernet and Wi-Fi, 1,492 on PPPoE.
MSS
Largest TCP payload a host will accept in one segment. MTU minus 40 (IPv4) or 60 (IPv6).
PMTU black hole
Big packets silently dropped because the ICMP that would explain why is filtered.
TUN / TAP
Virtual adapters that hand IP packets / Ethernet frames to a program instead of a wire.
Full / split tunnel
Everything through the VPN / only chosen prefixes through the VPN.
Cryptokey routing
WireGuard's AllowedIPs: one list that is both the tunnel's routing table and its access control.
Coordination server
Tailscale's control plane: distributes public keys, addresses and policy. Carries no traffic.
DERP
Tailscale's relays. Forward encrypted WireGuard packets over HTTPS when no direct path works.
MagicDNS
Tailscale's per-device resolver at 100.100.100.100 that names every machine in the tailnet.
Link monitor
The part of an overlay client that listens for OS network-change events and re-checks paths.
Where this comes from

Sources