Chapter 05 · Part II · Moving data

TCP, UDP & ports

IP gets a packet to the right computer, usually. It doesn't promise the packet arrives, arrives once, or arrives in order, and it doesn't say which program on that computer should get it. Two protocols sit on top to fill the gap. TCP builds a reliable, ordered stream out of unreliable packets. UDP adds almost nothing, on purpose.

65,536port numbers per protocol, 0 to 65,535: a 16-bit field in every TCP and UDP header
3 segmentsto open a TCP connection: SYN, SYN-ACK, ACK. One round trip before any data moves
8 vs 20 bytesUDP's whole header, against TCP's minimum header (up to 60 with options)

Common mix-up: a port isn't a physical socket or a door in a wall. It's a number in the packet header that tells the receiving operating system which program gets the data. "Opening a port" on a firewall means "allow packets with this number through," nothing more.

Who gets the packet

Ports, sockets and the five-tuple

Your laptop has one IP address on the network but dozens of programs talking at once: a browser with twenty tabs, a music app, a mail client, a backup agent. When a packet arrives, the operating system has to decide which one gets it. The IP header only says "this computer." The answer is the port number, a 16-bit field at the start of every TCP and UDP header. There are two in each packet: a source port and a destination port.

Servers listen on ports everyone agrees on, so clients know where to knock: 443 for HTTPS, 53 for DNS, 22 for SSH, 25 for mail between servers. Clients don't need a famous number. When your browser opens a connection, the operating system picks a spare temporary port, an ephemeral port, for the source, and the server sends its replies back to it.

A socket is the program's handle on one end of this: an address and port the program has claimed. A TCP connection is identified by five things together, the five-tuple: protocol, source address, source port, destination address, destination port. Change any one and it's a different connection. That's how a single web server on port 443 holds thousands of simultaneous connections: every client brings a different address-and-port pair.

(TCP, 192.168.1.57:52114 → 203.0.113.80:443)
One connection, named by its five-tuple. Open a second tab to the same site and only the source port changes.
RangeNameUsed for
0 – 1023System (well-known)Core services assigned by IANA: 22 SSH, 25 SMTP, 53 DNS, 80 HTTP, 123 NTP, 443 HTTPS. On Unix-like systems only privileged programs may listen here.
1024 – 49151User (registered)Services registered with IANA, such as 3389 Remote Desktop, 5432 PostgreSQL. Registration is a convention, not an enforcement.
49152 – 65535Dynamic (ephemeral)Never assigned. Picked on the fly for the client side of connections. Windows uses this range; Linux defaults to 32768–60999.
Go deeper: seeing your own sockets

On Windows, netstat -ano lists every socket with its local and remote address, state and owning process ID; PowerShell's Get-NetTCPConnection does the same with filtering. On Linux, ss -tunap; on macOS, netstat -an or lsof -i. A line in state LISTEN is a server waiting for connections; ESTABLISHED is a live conversation; TIME_WAIT is a recently closed one (more on that below).

Ports are also why NAT works. A home router rewrites the private source address to its public one, and if two devices happen to pick the same source port, it rewrites the port too, keeping a table so replies find their way back. That trick, and why it breaks incoming connections, is Chapter 07.

Opening a conversation

The three-way handshake

TCP, the Transmission Control Protocol, promises the receiving program a stream of bytes that arrives complete and in order, or a clear error. To keep that promise, both ends need shared bookkeeping before any data moves. They set it up with three segments:

  1. SYN (synchronize). The client says: "I'd like to talk. My byte counter starts at x."
  2. SYN-ACK. The server says: "Fine. I acknowledge your x; I expect byte x+1 next. My own counter starts at y."
  3. ACK. The client says: "Got it, I expect your byte y+1." Both sides are now ESTABLISHED.

Why three and not two? Each side has to announce its own starting number and hear the other side confirm it. The SYN-ACK does two jobs in one segment. The final ACK confirms the server's number; without it, the server couldn't tell a real client from a stale or forged SYN.

The starting numbers, the initial sequence numbers, are chosen unpredictably, not from zero. If they were guessable, an attacker who could send packets with a forged source address could inject data into someone else's connection. Random starting points also keep a delayed segment from an old connection on the same five-tuple from being mistaken for part of a new one.

The handshake has a cost: one full round trip before the first byte of real data. Between New York and London that's around 70 milliseconds, and an HTTPS page adds TLS's own negotiation on top. Shaving those round trips is a big part of why newer protocols exist.

Go deeper: flags and SYN floods

The TCP header carries a row of one-bit flags. The ones that matter here: SYN, ACK, FIN (I'm done sending), RST (abort, this connection doesn't exist), and PSH (deliver to the application now). A segment can carry several at once: SYN+ACK in the handshake, FIN+ACK when closing.

A server that receives a SYN has to remember it while it waits for the final ACK. A SYN flood sends huge numbers of SYNs from forged addresses and never completes them, filling that memory. The common defense, SYN cookies, encodes the needed state into the server's own initial sequence number, so the server can remember nothing until the final ACK proves the client is real.

Every byte counted

Sequence numbers, acks and retransmission

TCP numbers bytes, not packets. Every segment carries a sequence number: the position of its first byte in the stream. Segments flowing back carry an acknowledgment number: the next byte the receiver expects. Ack 7,301 means "I have everything up to byte 7,300, send me 7,301 onward." It's cumulative: one ack covers everything before it.

The receiver uses sequence numbers to put segments back in order and to throw away duplicates. The sender uses acks to know what got through. Anything sent but not yet acknowledged is kept in memory, ready to send again.

Lost segments get noticed two ways:

  • The retransmission timer. The sender keeps a running estimate of the round-trip time and how much it varies. If the oldest unacknowledged segment isn't acked within a timeout (the RTO) a bit longer than that, it's sent again, and the timeout doubles in case the network is struggling. The standard suggests at least one second; Linux uses a 200 millisecond floor.
  • Duplicate acks. If segment 3 of 6 is lost, the receiver gets 4, 5 and 6 but can't move its ack past 3. It keeps repeating "still waiting for 3." Three duplicate acks in a row is strong evidence of a loss rather than a reorder, so the sender resends segment 3 at once without waiting for the timer: fast retransmit. When 3 arrives, the receiver, which kept 4 to 6, jumps its ack straight to 7.

The program reading the stream never sees any of this. It just sees bytes arrive in order, occasionally after a pause. That pause is the price of reliability, and it's the problem for anything live: a voice packet resent 200 milliseconds late is worse than no packet.

ack = last in-order byte received + 1
Cumulative: segments after a gap are held by the receiver, but the ack can't move past the gap until it's filled.
Go deeper: how the timeout is computed

RFC 6298 keeps two running numbers per connection: SRTT, a smoothed round-trip time, and RTTVAR, how much it varies. Each new measurement R updates them: RTTVAR ← ¾ RTTVAR + ¼ |SRTT − R|, then SRTT ← ⅞ SRTT + ⅛ R. The timeout is RTO = SRTT + 4 × RTTVAR, never less than the floor. A connection with a steady 50 ms round trip ends up with an RTO just above the floor; a jittery Wi-Fi link gets a longer one.

Plain cumulative acks can only describe one hole at a time. The selective acknowledgment option (SACK, RFC 2018) lets the receiver list the blocks it holds past the gap, so the sender resends only what's missing. Nearly every modern TCP uses it.

How fast to send

Flow control and congestion control

Waiting for each segment's ack before sending the next would be painfully slow: one segment per round trip. So TCP keeps many segments in flight at once, up to a limit called the window. As acks come back, the window slides forward and more can be sent. Two separate limits set its size.

Flow control protects the receiver. Every ack carries a receive window: how many more bytes the receiver has buffer space for right now. If the reading program is slow, the window shrinks, and at zero the sender pauses. It's the receiver saying "slow down, I'm full."

Congestion control protects the network. The sender doesn't know how much capacity the path has, or how many other connections share it, so it keeps its own estimate, the congestion window, and probes:

  • Slow start. A new connection begins with a small window (ten segments on modern systems) and adds one segment for every ack, which doubles the window every round trip. "Slow" is ironic: it grows exponentially.
  • Congestion avoidance. Past a threshold, growth becomes gentle: about one extra segment per round trip.
  • Back off on loss. A lost segment is read as "the network is full." On fast retransmit the window is cut in half; on a timeout, it collapses to one segment and slow start begins again.

Grow by adding, shrink by halving: additive increase, multiplicative decrease, AIMD. Plotted over time it makes a sawtooth. Many connections all doing it share a bottleneck roughly fairly, and the internet doesn't collapse under its own traffic, which it did in 1986 before these rules existed. Linux, Windows and macOS now default to a variant called CUBIC that regrows faster on high-speed paths, but the shape is the same idea.

The window, not the link speed, often limits a single download. A connection can never move more than one window per round trip.

throughput ≤ window ÷ round-trip time
A 64 KB window over a 100 ms round trip caps out at 640 KB/s, about 5.2 Mbit/s, even on a gigabit line. Filling a 50 Mbit/s path at 100 ms needs 625 KB in flight: the bandwidth-delay product.
Go deeper: window scaling and the bandwidth-delay product

The receive window field in the TCP header is 16 bits, so on its own it maxes out at 65,535 bytes. That was plenty in 1981 and is far too small for a fast, long path today. The window scale option (RFC 7323), agreed during the handshake, multiplies the field by a power of two, allowing windows up to about a gigabyte.

The amount of data that has to be in flight to keep a path busy is its bandwidth-delay product: rate × round-trip time. 1 Gbit/s across a 70 ms transatlantic path is 8.75 MB in the air at once. If either window is smaller, the link sits partly idle no matter how fast it is. That's why a file copy over a VPN to a far-away site can crawl even when both ends have fast connections.

Loss-based congestion control has a known weakness: it fills the bottleneck's buffer before it sees a loss, so queues stay full and latency climbs. This is bufferbloat, and it's why a big upload can make a video call stutter on the same connection.

Instrument 1

Sliding-window simulator

A sender pushes 1,460-byte segments through a 50 Mbit/s bottleneck to a receiver. Set the window, the round-trip time and the loss rate. Orange boxes are data in flight, teal ticks are acks coming back, ✕ marks a loss. The strip shows the stream: acknowledged, in flight, waiting for retransmission, and the bracket is the window. Turn on congestion control to watch slow start and the AIMD sawtooth in the graph.

Goodput (last 10 RTTs)–
Ceiling: window ÷ RTT–
Window now–
Lost · retransmitted–

Data segments are lost at the rate you set; acks are not. The receiver holds out-of-order segments, and the sender uses fast retransmit, NewReno-style recovery and a timer with a 200 ms floor.

Saying goodbye

FIN, RST and the long wait in TIME-WAIT

A TCP connection is really two one-way streams, and each side closes its own. A side that has finished sending sends a FIN. The other side acks it, can keep sending for as long as it likes (that's a "half-close"), and eventually sends its own FIN, which gets acked in turn. Four segments, though the middle two are often combined into one.

The side that closed first doesn't forget the connection right away. It sits in TIME-WAIT for twice the maximum segment lifetime. That covers two risks: its final ACK might be lost (so it must be around to resend it when the other side repeats its FIN), and stray delayed segments from this connection must die out before the same five-tuple can be used again. The standard's segment lifetime is two minutes, making TIME-WAIT four; Linux uses a fixed 60 seconds and Windows two minutes. A busy client making thousands of short connections to the same server can run out of ephemeral ports because they're all parked in TIME-WAIT.

RST, reset, is the hard stop. A host sends it when a segment arrives for a connection that doesn't exist: a SYN to a port where nothing is listening gets a RST back, which your program reports as "connection refused." Programs can also abort a connection with RST, and middleboxes such as firewalls sometimes inject one to kill a connection they don't like. Compare that with a silent drop: if a firewall simply discards the SYN, the client just waits, retries, and eventually reports a timeout. "Refused" means you reached the machine. "Timed out" usually means you didn't.

StateMeans
LISTENServer socket waiting for a SYN.
SYN-SENTClient has sent SYN, waiting for SYN-ACK.
SYN-RECEIVEDServer has replied SYN-ACK, waiting for the final ACK.
ESTABLISHEDOpen. Data flows both ways.
FIN-WAIT-1 / 2We sent FIN; waiting for its ack (1), then for the other side's FIN (2).
CLOSE-WAITThe other side sent FIN; we haven't closed yet. Lots of these means a program forgot to close its sockets.
LAST-ACKWe sent our FIN after theirs; waiting for the final ack.
TIME-WAITClosed first; lingering 2 × MSL to catch strays.
CLOSEDNo connection.
Instrument 2

Handshake & teardown stepper

Step through a whole connection one segment at a time. Sequence and ack numbers are computed from random starting points, exactly as TCP does: SYN and FIN each use up one sequence number, data uses one per byte. Try the refused port and the lost SYN too.

Client state–
Server state–
This segment–

No frills, on purpose

UDP, and who chooses it

The User Datagram Protocol, defined in 1980 on three pages, adds just two things to IP: ports, and an optional checksum. Its header is 8 bytes: source port, destination port, length, checksum. No handshake, no sequence numbers, no acks, no retransmission, no windows. A program hands UDP a message, the message goes out as one packet, and it arrives once, or not at all, possibly out of order. Nobody tells the sender which.

That sounds like a worse TCP, but for a lot of traffic it's what you want:

  • Voice and video calls. Audio arrives in chunks every 20 milliseconds or so. A late chunk is useless; it's better to skip it and conceal the gap than to stall the whole call waiting for a resend. Calls use RTP over UDP and handle loss themselves.
  • DNS. One small question, one small answer. A TCP handshake would triple the time. If no answer comes, the resolver just asks again.
  • Games. The newest position update replaces the last one. Resending old positions would be worse than useless.
  • Tunnels like WireGuard (and Tailscale, which is built on it). The traffic inside the tunnel is often TCP already, with its own retransmission. Carrying TCP inside TCP stacks two sets of timers that fight each other when packets are lost. WireGuard sends each encrypted packet as one UDP datagram and lets the inner protocol do its job. UDP is also much easier to get through NAT, as Chapter 09 shows.
  • QUIC and HTTP/3. Rather than change TCP, which lives inside every operating system and many middleboxes, QUIC builds its own reliable, encrypted streams in the application on top of UDP. It combines the transport and encryption handshakes, and a lost packet stalls only the one stream it belonged to.

The catch: a UDP application has to do for itself whatever it needs. It has to pace itself so it doesn't flood the network, cope with loss, and keep NAT mappings alive with periodic packets, because a router can't see a FIN to know when a UDP "connection" is over. It just forgets the mapping after a quiet spell.

FeatureTCPUDP
Setup3-way handshake (1 round trip)None: first packet carries data
DeliveryReliable, ordered byte streamEach datagram once or not at all, any order
Message edgesNone: a stream of bytesKept: one send = one datagram
Loss handlingRetransmits; the stream stalls until filledUp to the application
Rate controlFlow and congestion control built inUp to the application
Header20–60 bytes8 bytes
Typical usesWeb, email, SSH, file transfer, RDP's main channelDNS, voice, video, games, WireGuard, QUIC
Go deeper: Remote Desktop uses both

Microsoft's Remote Desktop Protocol listens on TCP 3389 and, on modern versions, also UDP 3389. The session starts over TCP; if UDP gets through, RDP moves screen updates onto a UDP transport that handles loss in a smarter way for interactive graphics, and falls back to TCP if it can't. When RDP runs inside a tunnel such as Tailscale, all of that rides inside the tunnel's own UDP packets. If the tunnel can't find a direct path and falls back to a relay, every screen update makes a detour, which is exactly the case Chapter 09 and Chapter 13 take apart.

Cheat sheet

Terms from this chapter

Port
A 16-bit number in TCP and UDP headers that picks which program on a host gets the data.
Ephemeral port
A temporary source port the OS picks for the client side of a connection.
Socket
A program's handle on one endpoint: an address, a port and a protocol.
Five-tuple
Protocol + source address + source port + destination address + destination port. Names one connection.
SYN / ACK / FIN / RST
TCP flags: start, acknowledge, finished sending, abort.
Sequence number
The position in the byte stream of a segment's first byte.
Acknowledgment number
The next byte the receiver expects. Cumulative.
RTO
Retransmission timeout: how long to wait for an ack before resending.
Fast retransmit
Resending at once after three duplicate acks, without waiting for the timer.
Receive window
Flow control: how much more the receiver can buffer right now.
Congestion window
The sender's own estimate of how much the network can take.
Slow start / AIMD
Double per round trip at first; later add one per round trip and halve on loss.
Bandwidth-delay product
Rate × round-trip time: the data that must be in flight to keep a path full.
TIME-WAIT
The state a socket lingers in after closing first, for 2 × MSL.
Datagram
A self-contained UDP message: one send, one packet.
Where the facts come from

Sources

  1. RFC 9293, Transmission Control Protocol (TCP) (header, flags, three-way handshake, state machine, TIME-WAIT and MSL). rfc-editor.org/rfc/rfc9293
  2. RFC 768, User Datagram Protocol. rfc-editor.org/rfc/rfc768
  3. RFC 6335, IANA Procedures for the Management of the Service Name and Transport Protocol Port Number Registry (system, user and dynamic ranges); IANA's port registry. rfc6335 · iana.org port registry
  4. RFC 6528, Defending against Sequence Number Attacks (unpredictable initial sequence numbers). rfc-editor.org/rfc/rfc6528
  5. RFC 6298, Computing TCP's Retransmission Timer (SRTT, RTTVAR, the 1-second minimum and initial RTO). rfc-editor.org/rfc/rfc6298
  6. RFC 5681, TCP Congestion Control (slow start, congestion avoidance, fast retransmit); RFC 6582, The NewReno Modification to TCP's Fast Recovery. rfc5681 · rfc6582
  7. RFC 6928, Increasing TCP's Initial Window (ten segments); RFC 9438, CUBIC for Fast and Long-Distance Networks. rfc6928 · rfc9438
  8. RFC 7323, TCP Extensions for High Performance (window scaling); RFC 2018, TCP Selective Acknowledgment Options. rfc7323 · rfc2018
  9. RFC 8085, UDP Usage Guidelines (congestion control, keepalives through NAT); RFC 3550, RTP; RFC 9000, QUIC. rfc8085 · rfc3550 · rfc9000
  10. J. A. Donenfeld, WireGuard: Next Generation Kernel Network Tunnel (why WireGuard runs over UDP). wireguard.com/papers/wireguard.pdf
  11. V. Jacobson, Congestion Avoidance and Control, SIGCOMM 1988 (the 1986 congestion collapse and the origin of slow start).
  12. Microsoft Learn, Remote Desktop and Windows networking documentation (RDP on TCP and UDP 3389; Windows dynamic port range 49152–65535). learn.microsoft.com