Chapter 10 · Part IV · Routing

Routing tables & longest-prefix match

Every computer, phone and router on the internet carries a small list that answers one question for every packet: which way out? This chapter is about that list, the one rule that reads it, and what happens when a laptop has a dozen possible ways out and has to pick one.

33possible IPv4 prefix lengths, /0 to /32, that a lookup weighs against each other
0.0.0.0/0the default route: matches all 4,294,967,296 IPv4 addresses, and loses to everything else
64 · 128typical starting TTL on Linux and macOS · on Windows. Each router subtracts one

Common mix-up: the next hop's address never goes into the packet. The destination IP stays the same from your laptop to the server. At each step, the router only uses the next hop's address to find its hardware (MAC) address, then re-wraps the same packet in a new Ethernet frame addressed to that neighbor.

The list

Every device has a routing table

It's tempting to think routing is something only routers do. It isn't. The moment your laptop wants to send a packet, it has the same problem a backbone router has: there is a destination address, and there are several doors out of the machine. Which one? The answer comes from the routing table, a short list the operating system keeps in memory and consults for every single packet it sends.

Each line, or route, says: for destinations in this range, hand the packet to this neighbor, out of this interface, and here is how much I like this option. That's four columns:

  • Destination prefix. A range of addresses, written as an address and a length: 192.168.86.0/24 means "every address whose first 24 bits match 192.168.86". Chapter 3 covers how those bits work.
  • Next hop (or gateway). The neighbor that should get the packet next. If the destination is on a network the machine is plugged straight into, there is no middleman, and the route says on-link or directly connected.
  • Interface. Which network adapter to send it out of: Wi-Fi, Ethernet, a VPN's virtual adapter.
  • Metric. A cost. When two routes are otherwise equal, the lower number wins.

A typical home laptop's table is short. Joining Wi-Fi at 192.168.86.21 with a /24 mask gives it a connected route for free: 192.168.86.0/24, on-link, out of Wi-Fi. Anything in that range is a neighbor; the laptop can find its MAC address with ARP and deliver it directly. DHCP also hands it a router address, 192.168.86.1, which becomes the default route: 0.0.0.0/0 via 192.168.86.1. Add a loopback route for talking to itself, and that's the whole internet, as far as the laptop is concerned. Everything nearby goes direct; everything else goes to the router and becomes its problem.

A router's table is the same idea at a larger scale. A home router has two or three lines. A router at the edge of a big network might have a few hundred. A router in the internet's core, one with no default route at all because it has to know the way to everywhere, carries close to a million IPv4 prefixes and a couple of hundred thousand IPv6 ones, and looks them up for hundreds of millions of packets a second.

route = (prefix, next hop, interface, metric)
Every routing table in the world, from a phone to a core router, is a list of these. The rest of the chapter is about how one gets chosen.
Go deeper: the RIB, the FIB, and the hardware that does the lookup

Routers keep two tables. The RIB (routing information base) holds every route the router has heard about, from every source, including the losers. The FIB (forwarding information base) holds only the winners, already resolved down to an outgoing interface and a neighbor's MAC address, laid out for fast lookup. Routing protocols write to the RIB; packets are switched using the FIB.

On a PC the lookup happens in software, a few hundred nanoseconds per packet. Big routers do it in silicon. One classic approach is a TCAM, a memory that compares the destination against every stored prefix at once and returns the longest match in a single clock cycle. Others walk a compressed binary tree (a trie) of prefixes, one bit or a few bits per step, which is exactly the bit-by-bit picture in the instrument below.

IPv6 works identically, with 128-bit addresses and prefixes from /0 to /128. The default route is written ::/0.

The one rule

The longest prefix wins

Here's the puzzle: the ranges in a routing table overlap. The default route 0.0.0.0/0 contains every address there is. Your LAN route 192.168.86.0/24 is a small slice inside it. A VPN might add 10.0.0.0/8, and inside that, 10.20.0.0/16. So when a packet heads for 192.168.86.40, two routes match it. Which one applies?

The rule, stated in the IPv4 router requirements (RFC 1812) and used by every IP stack, is longest-prefix match: of all the routes whose prefix matches the destination, use the one with the longest prefix, the one that pins down the most bits. Longer prefix, smaller range, more specific knowledge. The /24 says "I know exactly where these 256 addresses live." The /0 says "I know nothing; try over there." Specific knowledge beats a shrug.

Checking whether a route matches is one line of bit math. Take the destination address, AND it with the route's mask (the mask has 1s for the prefix bits and 0s for the rest), and compare the result with the route's prefix. If they're equal, the route matches. Then keep the matching route with the most 1s in its mask.

match ⇔ (destination AND mask) = prefix
Among all matching routes, take the longest mask. Only if two matches have the same length does the metric get a vote.

That ordering surprises people. The metric is not a general "preference" knob. A /24 with a terrible metric still beats a /0 with a perfect one, because the comparison by length happens first and the metric is only a tie-breaker between routes of the same length. This is exactly how full-tunnel VPNs take over a laptop without deleting the existing default route: instead of adding a second 0.0.0.0/0, OpenVPN's classic trick installs 0.0.0.0/1 and 128.0.0.0/1. Together those two halves cover the whole address space, and each is one bit longer than /0, so they win every lookup no matter what the metrics say. When the VPN disconnects, it removes them and the original default is still there, untouched.

Longest-prefix match is also what makes the internet's table manageable. A provider that owns 198.51.100.0/22 can announce that single route instead of 1,024 individual addresses, and a customer who needs part of it routed elsewhere can announce a more specific /24, which automatically wins for its slice. That aggregation scheme is CIDR, from RFC 4632. The same property has a dark side, which Chapter 11 covers: whoever announces the most specific prefix wins, even if it isn't theirs.

Go deeper: equal-cost multipath, and why /32 routes exist

If two routes have the same prefix length and the same metric, many systems use both. That's ECMP, equal-cost multipath. To keep one conversation's packets in order, the router hashes each flow's addresses and ports and pins that flow to one path; different flows spread across all of them. Data centers lean on this heavily.

A /32 (or /128 in IPv6) is a route to a single address. Hosts install one for their own address, pointing at loopback, so they recognize packets meant for themselves. VPN clients install one for the VPN server, pointing out of the physical adapter, so the tunnel's own encrypted packets don't try to go through the tunnel. Without that one line, the tunnel would swallow its own transport. Chapter 12 lets you break it on purpose.

Instrument 1

Longest-prefix-match router

Type a destination address or pick one. The canvas compares it, bit by bit, against every route in the table: teal bits match, a red bit is the first mismatch, and grey bits are past the route's prefix length, so they don't matter. The winner is highlighted. Edit the table freely; prefixes like 10.0.0.0/8 or default both work.

PrefixNext hopInterfaceMetric
Winning route–
Sent to–
Routes that matched–
Why it won–

Who writes the table

Connected, static, and learned

Routes arrive from a handful of places, and it helps to know which is which when you're reading a table and wondering where a line came from.

  • Connected routes appear by themselves when an interface comes up with an address. Give an adapter 192.168.86.21/24 and the system adds 192.168.86.0/24 out of that adapter. Unplug the cable and the route disappears.
  • Static routes are typed in by a person or a script: route add on Windows, ip route add on Linux, a line in a router's config. They never change unless someone changes them, which is both their strength and their weakness.
  • DHCP-supplied routes. The default gateway handed out by DHCP becomes a default route. DHCP can also push extra static routes (option 121, classless static routes).
  • VPN and overlay routes. Connecting a VPN adds routes for whatever the tunnel serves: a corporate 10.0.0.0/8, the two halves of the internet for full tunnel, or, for Tailscale, the 100.64.0.0/10 range its devices live in.
  • Learned routes come from routing protocols such as OSPF and BGP, where routers tell each other what they can reach. That's the whole of Chapter 11.

Routers that hear about the same prefix from two sources need a way to decide which source to believe before they ever compare metrics, because an OSPF cost of 20 and a BGP path don't measure the same thing. Cisco calls this administrative distance; other vendors call it route preference. Lower is more trusted: connected 0, static 1, external BGP 20, OSPF 110, RIP 120, internal BGP 200. These numbers are vendor conventions, not a standard, but most vendors order the sources the same way.

Go deeper: recursive next hops, and why a gateway must be reachable

A route's next hop has to be reachable through some other route, usually a connected one. When you add "10.20.0.0/16 via 192.168.86.5", the system looks up 192.168.86.5 in its own table, finds it on the connected Wi-Fi network, and resolves the route to "out of Wi-Fi, to whatever MAC address 192.168.86.5 has". If no connected route covers the gateway, the route is useless; Windows refuses to add it, and Linux complains that the gateway is unreachable unless you mark it onlink.

BGP takes this further: its next hops are often routers several hops away, resolved through the interior routing protocol. That two-level lookup is called recursive resolution, and it's how a single change inside a network can re-steer thousands of internet routes at once.

Tie-breaks

Metrics, and the trouble with two defaults

Plug a laptop into Ethernet while it's on Wi-Fi, and it ends up with two default routes, both 0.0.0.0/0, both pointing at the same home router, one out of each adapter. Longest-prefix match can't separate them: they're the same length. So the metric decides.

Windows computes an interface metric for each adapter automatically from its link speed: the faster the link, the lower the number. A wired gigabit adapter usually ends up around 25, Wi-Fi somewhere in the 30s to 50s depending on the link rate it negotiated. Each route also has its own route metric (0 for a gateway handed out by DHCP, 256 for on-link routes), and the number route print shows is the two added together. That's why a LAN route on a gigabit adapter shows 281: 256 plus 25. With both adapters up, the Ethernet default at 25 beats the Wi-Fi default at 35, so traffic leaves by the cable. You can override the automatic value with Set-NetIPInterface -InterfaceMetric.

Linux puts a metric on each route; NetworkManager gives wired defaults 100, Wi-Fi 600, and VPNs 50, so the same preference falls out. macOS ranks network services by the order in System Settings and keeps one primary default route.

Two defaults are usually fine. They stop being fine in three ways. First, if both adapters sit on the same subnet, the system has two connected routes for that subnet too, and replies can leave by a different adapter than the request came in on. Second, a metric is a static guess based on link speed, not a measurement: a Wi-Fi adapter connected at a high rate can outrank a slower wired link that is actually more reliable. Third, every adapter that comes and goes rewrites the table. Each change can move traffic to a different interface mid-conversation, and every program that watches the network for changes, VPN clients especially, is told to re-examine everything.

effective metric = route metric + interface metric
The Windows rule. Lowest wins, but only among routes of the same prefix length.
Go deeper: strong and weak host models, and source address choice

Which source address does a packet carry when a machine has several? On Windows (Vista onward) the default is the strong host model: a packet must leave from the interface that owns its source address, so the routing decision also picks the source. Linux by default uses the weak host model: any interface can send with any of the host's addresses, and it will even answer ARP for one adapter's address on another, a behavior called ARP flux that confuses people with two adapters on one subnet. Linux can be pushed toward strict behavior with the arp_ignore, arp_announce and rp_filter settings, or with policy routing (ip rule), which keeps several tables and chooses one by source address.

IPv6 adds its own source-selection rules (RFC 6724), preferring, among other things, an address whose scope and prefix best match the destination.

Instrument 2

Too many adapters

A Windows laptop, modeled on a real one that ended up with a dozen network adapters. Add or remove adapters, pick a destination, and watch which way the packet actually leaves. The simulator builds the routing table from the adapters that are present, runs longest-prefix match, then breaks ties on metric. Turn on the flapping adapter and watch the table churn.

Leaves on–
Because of route–
Default routes–
Route-table changes–

Play with it and a pattern shows up. Adding adapters rarely breaks the choice outright; the lowest metric still wins. What it breaks is stability. Each virtual switch, bridge and dead VPN adapter is one more interface for the operating system to bring up, tear down and renumber, and each event is broadcast to every program that listens for network changes. An overlay VPN client listens for exactly that, because when the network changes, its direct paths to peers may have changed too. Feed it a steady trickle of changes and it spends its time re-probing, re-binding and rebuilding paths, and in the meantime traffic falls back to the slowest path it has, a relay. Chapter 12 follows that thread to the end.

When tables disagree

Hop by hop, and the TTL that ends loops

A packet's route isn't planned in advance. Each router along the way looks up the destination in its own table, picks a next hop, and forgets about it. There is no shared map and no record in the packet of where it has been. That makes routing robust and fast, and it also means that if two routers' tables disagree, a packet can be passed back and forth between them forever: router A thinks the way to 203.0.113.0/24 is through B, and B thinks it's through A.

The safety valve is a single byte in the IPv4 header, the Time To Live (RFC 791). The sender sets it, usually to 64 on Linux and macOS or 128 on Windows. Every router that forwards the packet subtracts one. A router that brings it to zero throws the packet away and sends the sender an ICMP Time Exceeded message (type 11). IPv6 renamed the field Hop Limit (RFC 8200), which is more honest: it was always a hop count, never a time. A looping packet therefore dies after at most 255 hops instead of circling until the link melts.

Traceroute turns this into a measuring tool. It sends a probe with TTL 1, which the first router discards, revealing itself in the Time Exceeded reply. Then TTL 2, which reveals the second router, and so on until the destination answers. Each line of traceroute output is one router admitting it just killed your packet. When you see the same two addresses alternating down the list, you're looking at a loop.

Loops are usually short-lived: they appear for a few seconds while dynamic routing protocols re-converge after a change, or permanently when someone adds a static route that points back the way the packet came. A default route on router A pointing at B, and a default on B pointing at A, is the classic.

TTLnext = TTL − 1; TTL = 0 → drop + ICMP Time Exceeded
The only per-hop memory a packet has. It bounds how long a mistake can live, it doesn't fix the mistake.
Go deeper: what else a router does to each packet

Forwarding an IPv4 packet is a short checklist: verify the header checksum, look up the destination, decrement the TTL, recompute the checksum (since the TTL changed), resolve the next hop's MAC address, and send it in a fresh frame. If the packet is bigger than the outgoing link's MTU, either fragment it or, if the Don't Fragment bit is set, drop it and send back an ICMP "fragmentation needed" message. IPv6 routers skip the checksum (the IPv6 header doesn't have one) and never fragment; only the sender may.

The Linux kernel guards against a different kind of loop with reverse path filtering: if a packet arrives on an interface that the host wouldn't use to reply to its source, it's dropped as probably spoofed. That's a feature until you have two adapters and asymmetric routes, when it quietly eats legitimate traffic.

Hands on

Reading route print and ip route

Here is a trimmed Windows table from a laptop on Wi-Fi at 192.168.86.21, with an Ethernet dock also connected. Run route print -4 in a terminal to see your own.

IPv4 Route Table
===========================================================================
Active Routes:
Network Destination        Netmask          Gateway       Interface  Metric
          0.0.0.0          0.0.0.0     192.168.86.1    192.168.86.22     25
          0.0.0.0          0.0.0.0     192.168.86.1    192.168.86.21     35
        127.0.0.0        255.0.0.0         On-link         127.0.0.1    331
     192.168.86.0    255.255.255.0         On-link     192.168.86.22    281
     192.168.86.0    255.255.255.0         On-link     192.168.86.21    291
    192.168.86.21  255.255.255.255         On-link     192.168.86.21    291
        224.0.0.0        240.0.0.0         On-link     192.168.86.21    291

Reading it: two defaults (the dock's at 25 wins), two connected routes for the same /24 (again the dock wins, 281 against 291), a /32 for the laptop's own Wi-Fi address, and a multicast range. Notice Windows writes the mask out longhand and lists the interface by its address, not its name. Get-NetRoute and Get-NetIPInterface in PowerShell give the same data with names and the two metrics separated.

The Linux equivalent, ip route, is terser:

default via 192.168.86.1 dev enp0s31f6 proto dhcp metric 100
default via 192.168.86.1 dev wlp2s0 proto dhcp metric 600
192.168.86.0/24 dev enp0s31f6 proto kernel scope link src 192.168.86.22 metric 100
192.168.86.0/24 dev wlp2s0 proto kernel scope link src 192.168.86.21 metric 600

Each line is prefix, then via (next hop) and dev (interface). proto says who added it: kernel for connected routes, dhcp, static, or a routing daemon. scope link means on-link. The best tool on Linux skips the reading entirely: ip route get 198.51.100.7 runs the real lookup and prints the route the kernel would use, including the source address. Windows has Find-NetRoute -RemoteIPAddress for the same job, and macOS has route get.

QuestionWindowsLinuxmacOS
Show the tableroute print -4ip routenetstat -rn -f inet
Which route for X?Find-NetRoute -RemoteIPAddress Xip route get Xroute get X
Interface metricsGet-NetIPInterfaceip route (metric field)networksetup -listnetworkserviceorder
Add a routeroute add 10.20.0.0 mask 255.255.0.0 192.168.86.5ip route add 10.20.0.0/16 via 192.168.86.5route add -net 10.20.0.0/16 192.168.86.5
Trace the pathtracert Xtraceroute X · mtr Xtraceroute X
Cheat sheet

Terms from this chapter

Routing table
The list of routes a device consults for every outgoing packet.
Prefix
An address range written as address/length. The length counts the fixed leading bits.
Next hop
The neighbor that should receive the packet next. Used to find a MAC address; never written into the packet.
Connected route
A route created automatically for the network an interface sits on. On-link: no gateway needed.
Default route
0.0.0.0/0 (or ::/0). Matches everything and loses to any more specific route.
Longest-prefix match
The selection rule: of the matching routes, use the one with the longest prefix.
Metric
A route's cost. Breaks ties between routes of the same prefix length. Lower wins.
Interface metric
Windows' per-adapter cost, set automatically from link speed and added to each route's metric.
Administrative distance
On routers, how much a route's source is trusted (connected, static, OSPF, BGP), checked before metrics.
ECMP
Equal-cost multipath: spreading flows across several routes that tie exactly.
TTL / Hop Limit
A counter each router decrements. At zero the packet is dropped and an ICMP Time Exceeded goes back.
Where this comes from

Sources