Chapter 13 · Part V · In the field

The troubleshooting toolkit

Every network problem looks like "it's slow" or "it's broken" from the outside. Underneath, it is almost always one specific layer that failed, and a handful of small tools can find which one. This chapter is the method, the tools, and one real case worked from symptom to fix.

169.254.x.xthe address a computer gives itself when no DHCP server answers: a lease that never arrived
1,472 byteslargest ping payload that fits a 1,500-byte Ethernet frame unfragmented (1,500 − 20 IP − 8 ICMP)
30 hopshow far tracert and traceroute go by default, three probes per hop

Common mix-up: a ping that gets no answer does not mean the machine is down. Plenty of hosts and firewalls simply drop ICMP echo requests; Windows Defender Firewall blocks inbound ping by default. "No reply" means "no reply," and you need a second tool to know why.

The method

Work from the bottom up

A network is a stack of layers, each trusting the one beneath it (Chapter 1). A web page needs a name lookup, which needs a route to a DNS server, which needs a gateway, which needs an IP address, which needs a working link. If any rung is missing, everything above it fails, and from the top every failure looks the same: the page doesn't load.

So don't start at the top. Start at the bottom and climb, asking one question per rung:

  1. Link. Is the adapter up, connected, and the one you think it is?
  2. IP configuration. Does it have a sensible address, mask, gateway and DNS server?
  3. Gateway. Can you reach the router on your own subnet?
  4. DNS. Do names turn into addresses?
  5. Path. Can packets cross the internet to the far end, and which way do they go?
  6. Application. Is the service itself listening, and is it taking the path you expect?

The first rung that fails is where the problem lives. Everything above it is collateral damage, and there's no point debugging the browser while the laptop has no address.

The second habit is to split the problem in half. Pinging an internet address by number, not by name, is the classic split: if a ping to a numeric address like 198.51.100.7 works but names fail, the whole network is fine except DNS. If the gateway answers but nothing past it does, the problem is upstream of your house. Each test that passes throws away half the suspects.

The third habit: write down what you see before you change anything. Networks heal themselves in confusing ways. A DHCP lease renews, a VPN reconnects, a cache expires, and the symptom vanishes without telling you why. Copy the output of each command as you go. The case study below only worked because the evidence was saved before the fixes went in.

Go deeper: why bottom-up beats intuition

Experienced engineers often jump straight to a hunch, and a good hunch is fast. But hunches are drawn from past failures, and the failure in front of you is frequently a new one. The bottom-up walk costs a minute or two and is guaranteed to end at the right layer. Use the hunch to pick which test to run first, then let the ladder confirm it.

Two failures cause most of the confusion. Partial failures, such as one DNS server of two being dead, or 2% packet loss, produce symptoms that come and go. Run tests several times and look at loss percentages, not just "it worked once." Layer leakage is when a lower layer works for small things and fails for big ones: a path MTU problem lets ping and the TCP handshake through but stalls the first full-size packet. The ladder still works, but you have to test with realistic traffic at the path rung.

Rungs 1 and 2

What does this machine think it has?

Start by asking the computer what it believes about itself. Every operating system has a command that lists its network adapters and the addresses on them:

  • Windows: ipconfig /all, or in PowerShell Get-NetAdapter (link state, speed, driver) and Get-NetIPConfiguration (addresses, gateway, DNS).
  • macOS: ifconfig, plus networksetup -listallhardwareports to map names like en0 to "Wi-Fi."
  • Linux: ip addr and ip link. The old ifconfig may not even be installed.

Read the output with four questions. Which adapter carries the traffic? Look for the one with a default gateway. Is the address plausible for this network? Your home router hands out something like 192.168.1.x; an address from a different range means you're on a different network than you think, or on none at all. Is the mask right? A /24 (255.255.255.0) is normal at home. Which DNS servers? If they belong to a VPN you closed an hour ago, you've found something.

The single most useful smell lives here. An address starting with 169.254 means the computer asked for a DHCP lease, heard nothing, and gave itself a link-local address so it could at least talk to neighbors (Chapter 4). It is not a working configuration. There's no gateway, and nothing beyond the local wire is reachable. The Wi-Fi icon may still say "connected," because the radio link is fine; the failure is one rung up.

Count the adapters, too. A laptop that has had VPN clients, virtualization software and docking stations installed over the years can carry a dozen virtual adapters, most of them dead. Each one is a candidate for routes, DNS servers and confusion, and software that watches network changes has to look at all of them.

One more local table is worth knowing: the neighbor cache, shown by arp -a (Windows, macOS) or ip neigh (Linux). It lists the MAC address learned for each local IP (Chapter 2). If your gateway's entry is missing or marked incomplete, the machine can't even find the router on the wire. If the gateway's MAC suddenly changes, either the router was replaced or something on the network is answering for it.

Go deeper: reading ipconfig /all line by line

Physical Address is the adapter's MAC. DHCP Enabled: Yes plus a Lease Obtained time means DHCP worked; if you see Autoconfiguration IPv4 Address instead of IPv4 Address, it didn't, and that's the 169.254 address. (Preferred) after an address means it passed duplicate-address checks; (Duplicate) means another device already claims it. Default Gateway empty on the adapter you're using means no route off the subnet. Media State: Media disconnected marks an adapter with no link at all, which on a laptop is usually a VPN adapter that isn't connected.

On Hyper-V machines, the real network address often isn't on the physical adapter at all. Creating an external virtual switch rebinds the physical NIC to the switch and gives the host a new adapter called vEthernet (name), which carries the IP address. That surprises people the first time: the Wi-Fi or Ethernet adapter shows no IPv4 address, yet the machine is online.

Two quick repairs live here, too: ipconfig /release then /renew asks DHCP again, and ipconfig /flushdns clears the resolver cache. Neither fixes the cause, but both help confirm it.

Rungs 3 and 4

Ping the gateway, then split on DNS

Ping sends an ICMP echo request and waits for an echo reply (RFC 792). It reports whether replies came back, how long each round trip took, and the TTL of the reply. On Windows it sends four by default; on macOS and Linux it runs until you press Ctrl+C, so add -c 4.

Ping the gateway first. On a healthy home network the router answers in a millisecond or two over Ethernet, a few milliseconds over Wi-Fi. Tens of milliseconds, or occasional losses, to a router in the same room points at Wi-Fi interference or an overloaded router, and nothing downstream will be better than that. No reply at all, from a router that normally answers, means you're not really on its network.

Then make the DNS split. Ping a well-known address by number, then the same kind of host by name. If the number works and the name fails, routing is fine and DNS is broken. Look up a name directly to see which server is answering:

  • nslookup example.com on any system prints the server it asked and what came back. Add a server to test a specific one: nslookup example.com 192.168.1.1.
  • dig example.com (macOS, Linux) gives the full answer section, the TTLs and the query time. dig @192.168.1.1 example.com asks one server.
  • Resolve-DnsName example.com is the PowerShell version and uses the same resolver path as Windows applications.

A timeout from the configured server, but a quick answer when you ask your router or a public resolver directly, means the configured server is wrong or unreachable. A common source is a VPN client that set its own DNS server when it connected and failed to put the old one back.

Reply from 192.168.1.1: bytes=32 time=2ms TTL=64
Who answered, how big the probe was, the round-trip time, and the TTL left on the reply. Replies usually start at 64 (Linux, macOS, most routers) or 128 (Windows), so a reply with TTL 57 has probably crossed 7 routers.
Go deeper: loss, jitter and what ping can't tell you

Ping measures the network and the target's willingness to answer. Routers treat ICMP addressed to themselves as low-priority work done by a busy control processor, so a router can show high or lost replies to ping while forwarding real traffic perfectly. Loss that appears at one hop and disappears at the next is almost always this, not a real problem.

Look at the spread, not just the average. Round trips of 2, 3, 2, 180 milliseconds to a nearby router suggest Wi-Fi retransmissions or a queue that fills up under load ("bufferbloat"). For a cleaner test, ping during a big download: if latency to the gateway jumps from 2 ms to 200 ms, the router's buffers are too deep.

Ping is ICMP. It doesn't prove that TCP port 443 or UDP port 3389 gets through a firewall. For that, use the application, or a port test such as Test-NetConnection host -Port 443 in PowerShell or nc -vz host 443 on macOS and Linux.

Rung 5

Traceroute, and how TTL makes it work

Every IPv4 packet carries an 8-bit Time to Live field (RFC 791); IPv6 calls it the Hop Limit (RFC 8200). Each router that forwards the packet subtracts one. A router that brings it to zero must throw the packet away, and it usually sends the sender an ICMP Time Exceeded message (type 11, code 0) from its own address. The rule exists so packets caught in a routing loop die instead of circling forever.

Traceroute turns that safety rule into a map. It sends a probe with TTL 1: the first router discards it and replies Time Exceeded, revealing its address and the round-trip time. Then TTL 2, which dies at the second router. Then 3. Each probe reaches one hop further, until a probe finally arrives at the destination, which answers in its own way, and the trace stops.

  • Windows tracert sends ICMP echo requests. The destination answers with an echo reply.
  • macOS and Linux traceroute send UDP datagrams by default, to high ports starting at 33434 that nothing listens on. The destination answers with ICMP Port Unreachable (type 3, code 3). traceroute -I switches to ICMP, and on Linux -T uses TCP SYNs, which pass firewalls that block the others.

Both send three probes per hop and give up after 30 hops. Add -d (Windows) or -n (macOS, Linux) to skip the reverse DNS lookup of every hop, which makes the trace much faster.

Reading the result: each line is a hop, with three round-trip times. Times should grow roughly with distance. A row of asterisks means no reply arrived for that TTL; a router that forwards fine but doesn't send Time Exceeded, or rate-limits it, shows up as stars and then the trace carries on past it. Stars all the way to hop 30 mean the destination itself, or a firewall in front of it, isn't answering the probe type you used.

The first few hops tell you about your own house. One private address (192.168.x.x, 10.x.x.x or 172.16–31.x.x) and then your ISP is normal: one router doing NAT. Two private hops in a row, say 192.168.86.1 and then 192.168.1.1, means a router behind a router: double NAT (Chapter 8). A hop in 100.64.0.0/10 is your ISP's carrier-grade NAT.

Go deeper: pathping, mtr and the hop that lies

A single trace is a snapshot. pathping (Windows) first traces the route, then pings every hop 100 times, 250 milliseconds apart, about 25 seconds per hop, and reports loss per router and per link. mtr (macOS via Homebrew, Linux) does the same continuously, refreshing a table of loss and latency for each hop. For intermittent problems, these are the right tools: run one for a few minutes while the problem happens.

The rule for reading them: loss only matters if it continues to the end. If hop 4 shows 40% loss but hops 5 to 9 show 0%, hop 4 is simply deprioritizing replies to you; packets passing through it are fine. If loss appears at hop 4 and every hop after it shows similar loss, the problem is at or just before hop 4.

Traceroute can also mislead about the path. Routes can be asymmetric: the Time Exceeded reply may come back by a different path than the probe took, so a latency jump at hop 6 may be about the return route. Load balancers can send each probe down a different parallel path, producing a trace that zigzags between routers. Tools like Paris traceroute keep the probe fields constant to stay on one path.

Finally, the address shown for a hop is the interface the router chose to send the Time Exceeded from, usually the one facing you, not necessarily the one your traffic left by.

Instrument 1

Traceroute simulator

Pick a network, then send probes one at a time. Each probe leaves with the TTL shown; every router it crosses takes one off, and the router that drops it to zero sends back Time Exceeded. One router in each path forwards traffic but never replies. Switch between the Windows and Linux styles to see the destination answer differently.

Last probe–
Answered by–
Reply–
NAT layers seen–

      

Addresses are from documentation and private ranges. Round-trip times are computed from each link's delay plus a little random jitter, which is why the three probes per hop differ.

Which door?

Route tables and the two-default-routes trap

Before a packet leaves, your own machine decides which adapter and which gateway to use by longest-prefix match on its routing table (Chapter 10). You can see the table:

  • Windows: route print, or Get-NetRoute. The Find-NetRoute -RemoteIPAddress 198.51.100.7 cmdlet tells you which route a specific destination would use.
  • macOS: netstat -rn; route -n get 198.51.100.7 answers the same "which way?" question.
  • Linux: ip route, and ip route get 198.51.100.7.

The line to find is the default route, 0.0.0.0/0 (shown as destination 0.0.0.0, mask 0.0.0.0 on Windows, or default elsewhere). It matches everything, so it's what carries all your internet traffic. There should normally be exactly one active default. When there are two, on two adapters, the system picks by metric: on Windows the effective cost is the route metric plus the interface metric, and the lowest wins. That's how a VPN takes over all your traffic, and how a half-dead VPN can keep holding it after the tunnel has stopped working.

Specific routes beat the default regardless of metric, because longer prefixes always win. A split-tunnel VPN adds routes like 10.0.0.0/8 through the tunnel and leaves the default alone (Chapter 12). If a VPN adds a route that overlaps your home subnet, your printer can vanish the moment you connect.

To see what's actually using the network, list connections: netstat -ano on Windows (the last column is the process ID), ss -tunap on Linux, netstat -an or lsof -i on macOS. A connection stuck in SYN_SENT means your machine sent the opening TCP handshake and heard nothing back, usually a firewall silently dropping it. Lots of ESTABLISHED connections to a relay you didn't expect is a clue of a different kind, as the case study shows.

Go deeper: interface metrics on Windows

Windows assigns each interface an automatic metric based on link speed: faster links get lower numbers. A 1 Gbit/s Ethernet adapter typically lands below a Wi-Fi adapter, so plugging in a cable quietly moves the default route to the wire. VPN clients often set their interface metric very low on purpose so that their route wins. Get-NetIPInterface shows the metric of every interface; Set-NetIPInterface -InterfaceAlias "Wi-Fi" -InterfaceMetric 25 changes one.

Route metrics also decide which DNS servers Windows tries first, because DNS servers are attached to interfaces. A dead VPN adapter with a low metric and a stale DNS server can make every name lookup wait for a timeout before falling back.

Pattern recognition

Five smells worth memorizing

After enough incidents, certain readings jump off the screen. These five explain a large share of home and small-office problems.

SmellWhere you see itWhat it meansUsual fix
169.254.x.xipconfig, ip addrNo DHCP lease. The machine self-assigned a link-local address; no gateway.Check the link, router's DHCP server and pool; renew the lease.
Gateway OK, no DNSping by IP works, by name fails; nslookup times outRouting works; the configured resolver is wrong, unreachable or down.Point DNS at the router or a working resolver; remove stale VPN DNS.
Two default routesroute print, ip routeTwo adapters both claim the internet; metrics decide which wins, and it may be the wrong one.Remove or reprioritize the stale one; fix the VPN's cleanup.
Double NATtwo private hops in tracert; private "external" address from port mappingA router behind a router. Inbound connections and port mapping only reach the inner one.Bridge mode on one router, or DMZ the inner router on the outer.
MTU black holesmall pings pass, ping -f -l 1472 times out; pages half-loadA link on the path has a smaller MTU and the "fragmentation needed" message is being dropped.Lower the MTU or clamp TCP MSS on the router; stop blocking ICMP.

The MTU black hole deserves a word because it's the strangest. Path MTU discovery (RFC 1191) sends full-size packets with the Don't Fragment bit set, and relies on any router with a smaller link to send back ICMP "fragmentation needed" (type 3, code 4) with the right size. If a firewall blocks that ICMP, the sender never learns. The TCP handshake, made of tiny packets, succeeds; the first full-size data packet vanishes, and the connection hangs. PPPoE links (MTU 1492) and tunnels (WireGuard's default is 1420) are the usual narrow spots.

To test, send unfragmentable pings of a chosen size. On Windows, ping -f -l 1472 host; on macOS, ping -D -s 1472 host; on Linux, ping -M do -s 1472 host. A payload of 1,472 plus the 8-byte ICMP header and 20-byte IP header is exactly 1,500. If that fails but 1,464 works, the path MTU is 1,492.

Go deeper: why double NAT hurts more than it looks

Outbound browsing works fine through two NATs; each just rewrites the source address and keeps a table entry (Chapter 7). The pain is inbound and peer-to-peer. Port-mapping protocols such as UPnP and NAT-PMP only talk to the router directly above you, so the inner router happily opens a port on its outside address, which is itself a private address on the outer router's network. Nothing from the internet can reach it. The tell-tale sign: a port-mapping client reports an "external address" in a private range.

Peer-to-peer tools then fall back on hole punching (Chapter 9), which has to get through both NATs at once. It often succeeds if both NATs keep a consistent mapping per inside port, and fails if either one picks a new outside port per destination. When it fails, the tool uses a relay, which works but adds latency and caps throughput.

Overlay networks

Tailscale's own diagnostics

Overlay VPNs add a layer on top of all this, and they bring their own tools. Tailscale connects devices with WireGuard tunnels, trying hard to make each tunnel direct, and falling back to its DERP relay servers when it can't (Chapter 12). Three commands tell you which you've got:

  • tailscale status lists every peer with its tailnet address and, for active peers, how it's connected: direct 203.0.113.45:41641 (a real UDP path) or relay "nyc" (bouncing through the New York DERP server).
  • tailscale netcheck examines your local network: whether UDP works, the public address and port that STUN servers see you from, whether that mapping changes with the destination, which port-mapping protocols your router speaks (UPnP, NAT-PMP, PCP), and the latency to each DERP region.
  • tailscale ping <peer> pings at the Tailscale level and reports the path each reply took, such as via DERP(nyc) or via 203.0.113.45:41641. It keeps trying for a direct path, so watching it switch from DERP to direct is the clearest proof a fix worked.

The key netcheck line is MappingVariesByDestIP. If it's true, your NAT gives a different outside port for every destination, which makes hole punching hard, and relays become likely. If it's false and you're still relayed, look for something that keeps resetting the connection or a second NAT the port mapping can't reach.

Go deeper: why relayed feels so much worse than the numbers say

DERP carries traffic over HTTPS, which means TCP. Your remote desktop session was already sending its own packets, often UDP, inside WireGuard; on a relayed path they ride inside a TCP stream. One lost packet on the TCP leg stalls everything behind it until it's retransmitted (head-of-line blocking). A relay also adds the distance to the relay and back, and DERP servers are shared and rate-limited, so throughput is capped. Remote desktop, which sends a burst of screen updates every time something moves, is exactly the kind of traffic that feels this as lag and smearing.

Case study

A laptop's remote desktop got laggy

The symptom was plain: Remote Desktop into a Windows laptop over Tailscale had become sluggish. Typing echoed late, windows smeared when dragged, and it was worse than it had been a few weeks earlier. Nothing was "down." Here is the ladder, rung by rung, with what each tool showed.

  1. Link: Get-NetAdapter. Twelve adapters. The real LAN connection wasn't on the Ethernet or Wi-Fi adapter at all: a Hyper-V external virtual switch had claimed the physical NIC, and the address lived on vEthernet (External). Alongside it sat a Windows Network Bridge, a vEthernet (Default Switch), the Tailscale adapter, and a museum of VPN drivers: a NordVPN TAP adapter, an OpenVPN Data Channel Offload adapter, a Sophos TAP adapter and a Cisco AnyConnect adapter, none of them in use.

    PS> Get-NetAdapter | Sort Status | Format-Table Name, InterfaceDescription, Status
    Name                   InterfaceDescription                         Status
    ----                   --------------------                         ------
    vEthernet (External)   Hyper-V Virtual Ethernet Adapter             Up
    Ethernet 2             Realtek USB GbE Family Controller            Up      # bound to the vSwitch, no IP
    Wi-Fi                  Intel(R) Wi-Fi 6E AX211 160MHz               Up      # member of the bridge
    Network Bridge         Microsoft MAC Bridge Miniport                Up
    vEthernet (Default...  Hyper-V Virtual Ethernet Adapter #2          Up
    Tailscale              Tailscale Tunnel                             Up
    NordVPN TAP            TAP-NordVPN Windows Adapter V9               Disconnected
    OpenVPN DCO            OpenVPN Data Channel Offload                 Disconnected
    Sophos TAP             TAP-Windows Adapter V9                       Disconnected
    Cisco AnyConnect       Cisco AnyConnect Virtual Miniport Adapter    Disabled
    Bluetooth Network...   Bluetooth Device (Personal Area Network)     Disconnected
    vEthernet (WSL)        Hyper-V Virtual Ethernet Adapter #3          Up

    Adapter names and descriptions here are representative of what Windows prints for these drivers; the shape of the list is what matters.

  2. IP config: ipconfig /all. vEthernet (External) had 192.168.86.120/24, gateway 192.168.86.1, DNS 192.168.86.1, a valid DHCP lease. Rungs 1 and 2 pass, but with a note: that's a lot of adapters coming and going for one laptop.

  3. Gateway and DNS: ping 192.168.86.1, nslookup. 1–3 ms, no loss; names resolved instantly. The local network is healthy. Whatever is wrong is further up.

  4. Path: tracert -d. The first two hops gave it away:

    C:\> tracert -d 198.51.100.7
    Tracing route to 198.51.100.7 over a maximum of 30 hops
    
      1     2 ms     1 ms     2 ms  192.168.86.1
      2     3 ms     3 ms     4 ms  192.168.1.1
      3     9 ms     8 ms     9 ms  198.51.100.1
      ...

    Two private routers in a row: the laptop's router (192.168.86.1) is plugged into another router (192.168.1.1), probably the ISP's gateway. Double NAT.

  5. Path, Tailscale's view: tailscale netcheck. UDP worked and the mapping didn't vary by destination, both good signs. The router answered NAT-PMP port-mapping requests. But the port mapping it granted named an external address of 192.168.1.222: the inner router's own address on the outer router's network. A private "external" address is the signature of double NAT. The public address the STUN servers saw was a different one entirely.

    C:\> tailscale netcheck
    Report:
    	* UDP: true
    	* IPv4: yes, 203.0.113.45:41641        # what the internet sees
    	* IPv6: no, but OS has support
    	* MappingVariesByDestIP: false
    	* PortMapping: NAT-PMP
    	* Nearest DERP: New York City
    	* DERP latency:
    		- nyc: 11.4ms  (New York City)
    		- ...
    # daemon log, paraphrased: NAT-PMP mapping granted, external address 192.168.1.222
  6. Application: tailscale status and tailscale ping. The RDP peer was active; relay "nyc". Every screen update was making a round trip through a relay server instead of going straight across. Pinging it confirmed the path, and the "direct connection not established" message after the run is the giveaway:

    C:\> tailscale ping client
    pong from client (100.64.0.12) via DERP(nyc) in 41ms
    pong from client (100.64.0.12) via DERP(nyc) in 38ms
    pong from client (100.64.0.12) via DERP(nyc) in 44ms
    ...
    direct connection not established
  7. The daemon itself. Task Manager showed tailscaled using noticeable CPU on an idle machine, and its log was full of link-change events. Every time one of those virtual adapters changed state, the daemon re-examined the network: re-ran its checks, rebuilt its list of candidate endpoints, and restarted path discovery. With twelve adapters, a bridge and two virtual switches, it was never quiet long enough for a direct path to settle.

The diagnosis. Two problems compounding. First, double NAT: the inner router could map a port, but only onto a private address, so nothing outside could use it, and hole punching had to get through two NATs. Second, adapter churn: the overlay kept tearing down its own path discovery whenever a dead VPN adapter or the bridge twitched. Either alone might have been survivable. Together, the connection lived on the relay.

The fix list, in the order it was done:

  1. Remove dead adapters. Uninstall the NordVPN, OpenVPN, Sophos and Cisco clients that weren't in use, which removes their virtual adapters. Disabling them in Network Connections is a quick first step; uninstalling is permanent.
  2. Drop the bridge. The Hyper-V external switch already connects the virtual machines to the LAN. The Network Bridge was a second, overlapping way of doing the same thing. Delete it.
  3. Remove one NAT layer. Put the outer router (the ISP's gateway) into bridge mode so the inner router gets the public address. If the ISP box can't bridge, set its DMZ host to the inner router's outside address, 192.168.1.222, so every unsolicited inbound packet is passed straight through. Either way, NAT-PMP mappings on the inner router become reachable from the internet.
  4. Reboot. Removing network drivers and bridges leaves changes pending until restart. Reboot once, not three times in the middle.
  5. Confirm. Run tailscale netcheck again (the mapping's external address should now be public), then tailscale ping until it reports a direct path.
C:\> tailscale ping client
pong from client (100.64.0.12) via DERP(nyc) in 40ms
pong from client (100.64.0.12) via 203.0.113.45:41641 in 9ms

The first reply still came through the relay while the direct path was being negotiated; the second arrived direct. (If the client sits on the same LAN, the direct address shown will be the LAN one, 192.168.86.120:41641, and the time a couple of milliseconds.) tailscale status now showed direct, the daemon went quiet, and the remote desktop felt local again.

Go deeper: what was certain, and what was inference

Certain from the evidence: twelve adapters, a bridge and an external vSwitch; the LAN address on the virtual adapter; two private hops in the trace; a port mapping with a private external address; the peer on a DERP relay; direct after the fixes.

Inference: exactly how much each cause contributed. Tailscale can often punch through double NAT on its own when both routers keep stable mappings, and plenty of machines run with several VPN adapters installed. The ordering of the fixes, cheapest and most certain first, means the case doesn't prove which one alone would have been enough. That's normal in the field. The aim is a network that's simpler and a symptom that's gone, with the evidence written down for next time.

Instrument 2

Diagnosis game

Four broken networks. For each, read the symptom, pick a tool, read its output, and pick the next. The ladder above the console lights up as you test each rung. When you think you know, name the cause. Fewer tools is better, but the right answer matters most.


      

Name the cause
Tools used0
Rungs tested0 of 6
Verdict–

Cheat sheet

The same job on three systems

JobWindowsmacOSLinux
Addressesipconfig /allifconfigip addr
AdaptersGet-NetAdapternetworksetup -listallhardwareportsip link
Routesroute printnetstat -rnip route
Which route?Find-NetRoute -RemoteIPAddressroute -n getip route get
Neighborsarp -aarp -aip neigh
Pingping hostping -c 4 hostping -c 4 host
MTU testping -f -l 1472ping -D -s 1472ping -M do -s 1472
Tracetracert -dtraceroute -ntraceroute -n
Trace + losspathpingmtrmtr
DNS lookupnslookup, Resolve-DnsNamedigdig
DNS cacheipconfig /flushdnssudo dscacheutil -flushcacheresolvectl flush-caches
Connectionsnetstat -anonetstat -an, lsof -iss -tunap
Overlaytailscale status · tailscale netcheck · tailscale ping <peer> — the same everywhere

On macOS, flushing the cache fully also needs sudo killall -HUP mDNSResponder. resolvectl applies to Linux systems using systemd-resolved. mtr usually needs installing.

Cheat sheet

Terms from this chapter

ICMP
The Internet Control Message Protocol: the network's error and diagnostic messages, including echo (ping), Time Exceeded and Destination Unreachable.
TTL / Hop Limit
A counter in every IP packet that each router decrements. At zero the packet is dropped and the sender is told.
Time Exceeded
ICMP type 11. The message a router sends when it drops a packet whose TTL ran out. Traceroute is built on it.
Round-trip time
How long a probe takes to get there and back. What ping and traceroute report in milliseconds.
Link-local (169.254/16)
Self-assigned addresses used when DHCP fails. Valid only on the local wire.
Default route
0.0.0.0/0, the route that matches everything. Normally one; two is trouble.
Path MTU
The smallest MTU of any link between two hosts. Packets bigger than it must be fragmented or dropped.
MTU black hole
A path where oversized packets are silently dropped because the ICMP error that should report it is blocked.
Double NAT
Two layers of address translation, usually a router plugged into another router.
DERP
Tailscale's relay servers. Always work, cost latency and throughput.
Sources

Where these facts come from

  1. RFC 792 — Internet Control Message Protocol (echo, Time Exceeded, Destination Unreachable codes).
  2. RFC 791 — Internet Protocol (the TTL field).
  3. RFC 8200 — IPv6 (Hop Limit).
  4. RFC 1122 — Requirements for Internet Hosts.
  5. RFC 1191 — Path MTU Discovery; RFC 4821 — Packetization Layer Path MTU Discovery.
  6. RFC 3927 — Dynamic Configuration of IPv4 Link-Local Addresses (169.254/16).
  7. RFC 2131 — DHCP; RFC 1034 / 1035 — DNS.
  8. RFC 1918 — private address ranges; RFC 6598 — shared address space 100.64.0.0/10; RFC 5737 — documentation addresses.
  9. RFC 4787 — NAT behavioral requirements for UDP; RFC 6886 — NAT-PMP; RFC 5389 — STUN.
  10. Microsoft Learn, Windows commands: tracert, pathping, ipconfig, netstat.
  11. Microsoft Learn: Get-NetAdapter; Create a virtual switch for Hyper-V.
  12. Tailscale: CLI reference (status, netcheck, ping); DERP servers; How NAT traversal works.
  13. Linux man pages: ip-route(8), ss(8).