← all field notes
essay / August 22, 2026•30 min read

TCP from the inside

tcpnetworkinglinuxinternals

Everyone has debugged a TCP problem. Almost nobody can explain one.

The page fires at 2 a.m. p99 latency has gone from 4 ms to 4 seconds. CPU is flat, the database is bored, the application logs say nothing at all. You SSH in, run ss -s, and find forty-seven thousand sockets sitting in TIME_WAIT.

Most engineers restart the service at this point, watch the graph recover, and file it under mysteries. It works, because a restart wipes the socket table. It also guarantees you will be back here in a week, because nothing you did addressed why forty-seven thousand connections were being created and destroyed in the first place.

TCP is the most read-about and least understood protocol in production systems. Everyone can draw SYN, SYN-ACK, ACK on a whiteboard. Far fewer can say what number is in the ACK field and why it is that number, what happens to the connection when a single segment is dropped, or which of the two windows — the one the receiver advertises or the one the sender invented — is actually limiting their throughput right now.

This post is the layer underneath. Seven things that determine how TCP behaves in production: the handshake, reliable delivery, flow control, congestion control, the state machine, what the Linux kernel actually runs, and the gotchas that page you. Each one comes with a visualizer you can break on purpose.

Start with the three packets everyone thinks they know.

↳ the handshake, packet by packet

CLOSED
client state
LISTEN
server state
0/3
segments on the wire
clientserver:54312:443nothing on the wire yet — press “send next segment”

idle

The client socket is CLOSED, the server socket is in LISTEN. Nothing exists yet — no state is allocated on the server until the SYN arrives.

client ISN

1320745829

server ISN

3811002447

Part 1: The handshake is a state exchange, not a greeting

The three-way handshake is usually taught as a politeness ritual — hello, hello back, thanks. That framing hides the actual work. The handshake exists to do one thing: make both ends agree on where each direction’s byte stream starts. A TCP connection is two independent streams, one in each direction, and each one needs its own starting sequence number.

So the client picks an Initial Sequence Number and puts it in the SYN. The server picks its own, entirely unrelated ISN and puts that in the SYN-ACK — while separately acknowledging the client’s. The third packet closes the loop. Three packets, because there are two ISNs to communicate and each one needs acknowledging, and one of the four logical messages can be piggybacked.

Why the ISN is random

It would be much simpler to start every connection at zero. It would also be a security hole and a correctness bug at the same time.

The correctness half: IP networks reorder and delay. A segment from a previous connection on the same four-tuple can turn up minutes later. If sequence numbers always started at zero, that straggler would look like perfectly valid data for the current connection, and TCP would deliver it to your application. Random ISNs make the odds of an old segment falling inside the current window vanishingly small.

The security half: if the ISN is predictable, an off-path attacker who can guess it can inject data into your connection, or forge an RST and kill it, without ever seeing a packet. Linux generates ISNs from a keyed hash over the four-tuple plus a clock, per RFC 6528, so they are unpredictable to anyone who does not already know the tuple.

Why three packets and not two

The classic answer is “to prevent half-open connections,” which is correct and unilluminating. The concrete failure looks like this. A client sends a SYN. The network stalls it. The client times out and retries on a new connection, gets served, and goes away. Twenty seconds later the original SYN finally arrives at the server.

With a two-way handshake, the server would allocate a connection, reply, and consider the connection open. There is no third message, so there is no opportunity for the client to say “I never asked for this.” The server sits there with a live socket, a receive buffer, and possibly a worker thread, serving nobody. With three, the client receives a SYN-ACK for a connection it does not have, replies with an RST, and the server tears the whole thing down.

The teardown needs four, for the opposite reason

Opening is symmetric — neither side can send until both are ready. Closing is not. A TCP connection is two simplex streams, and they are shut down independently. When your side calls close(), you are saying “I will send no more bytes.” You are not saying anything about the peer, and it may have plenty left to send.

So each direction gets its own FIN and its own ACK: four segments. The state between them — one side closed, the other still sending — is called a half-close, and it is a real, legitimate, useful state. It is what shutdown(fd, SHUT_WR) gives you deliberately.

Both SYN and FIN consume a sequence number even though neither carries payload. That is why every ACK in the visualizer above comes back as ISN+1, and why the final ACK of the teardown is ISN+2. It is also what makes them reliable: because they occupy sequence space, a lost FIN gets retransmitted like any other unacknowledged data.

$ tcpdump -n -i any 'tcp port 443' -c 6

IP 10.0.1.7.54312 > 10.0.4.9.443: Flags [S],  seq 1320745829,                  win 64240, options [mss 1460,sackOK,TS,nop,wscale 7]
IP 10.0.4.9.443 > 10.0.1.7.54312: Flags [S.], seq 3811002447, ack 1320745830,  win 65160, options [mss 1460,sackOK,TS,nop,wscale 7]
IP 10.0.1.7.54312 > 10.0.4.9.443: Flags [.],               ack 3811002448,     win 502
IP 10.0.1.7.54312 > 10.0.4.9.443: Flags [P.], seq 1:518,   ack 1,              win 502, length 517
IP 10.0.4.9.443 > 10.0.1.7.54312: Flags [.],               ack 518,            win 509
IP 10.0.4.9.443 > 10.0.1.7.54312: Flags [P.], seq 1:2921,  ack 518,            win 509, length 2920

Two things in that capture are worth naming. The options in the SYN — mss, sackOK, wscale 7 — are negotiated once, in the handshake, and never again. Window scaling in particular is all-or-nothing for the life of the connection, which is why an old middlebox that strips the option can cap a modern connection at 64 KB in flight forever. And notice that tcpdump switches to relative sequence numbers after the handshake: seq 1:518 means bytes 1 through 517 of this direction’s stream. That relabelling is a display convenience, and it is also the correct mental model.

Part 2: Reliable delivery is a numbering scheme and a timer

TCP runs on top of IP, which offers no guarantees whatsoever. Packets may be dropped, duplicated, reordered, or corrupted. Every reliability property you get from TCP is built out of exactly two mechanisms: numbering every byte, and retransmitting anything not acknowledged in time. That is the whole thing.

Sequence numbers count bytes, not packets

This is the most common misreading, and it changes how everything else works. A sequence number is a byte offset into that direction’s stream. If you send a 1460-byte segment starting at sequence 5000, the next segment starts at 6460. Nothing counts packets anywhere in the protocol.

That is why TCP is a stream and not a message protocol. Your three 100-byte write() calls may arrive as one 300-byte read(), or as a 250-byte read followed by a 50-byte read, and both are correct behaviour. Every framing bug in every hand-rolled protocol comes from expecting otherwise.

An ACK is cumulative, and that makes lost ACKs cheap

ack=7000 does not mean “I got the segment starting at 7000.” It means “I have every byte below 7000, and 7000 is what I want next.” It is a high-water mark, not a receipt.

The consequence is that losing an ACK usually costs nothing at all. If the ACK for byte 7000 is dropped but the ACK for byte 10000 arrives, the second one already tells the sender everything the first one would have. Only the last ACK of a burst is genuinely load-bearing, and if that one is lost the RTO cleans it up.

Losing a segment is a different matter, and the visualizer below lets you do both so you can watch the asymmetry. Drop a segment and everything behind it stacks up in the receiver’s out-of-order queue while it repeats the same ACK number. Drop an ACK and, almost always, nothing happens.

↳ loss, cumulative ACKs and the retransmit timer

—
SRTT
—
RTTVAR
1000ms
RTO
0
duplicate ACKs

RTO = SRTT + 4·RTTVAR, floored at 200ms (Linux TCP_RTO_MIN) and doubled on every timeout.

sender — send buffer, snd_una = 1000

seq 1000queued
seq 2460queued
seq 3920queued
seq 5380queued
seq 6840queued
seq 8300queued
seq 9760queued
seq 11220queued
seq 12680queued
seq 14140queued
data →← ACKs

receiver — rcv_nxt = 1000

—
—
—
—
—
—
—
—
—
—
retransmit timer on seg 0idle
press play — the sender opens a 4-segment window and walks it

The retransmission timeout adapts, because the network moves

How long should a sender wait before deciding a segment is lost? Too short and you flood a healthy network with duplicates. Too long and every loss costs you seconds. And the right answer on a 0.3 ms path inside a rack is four orders of magnitude away from the right answer on a satellite link.

So TCP measures. Every unambiguous ACK produces an RTT sample, and Jacobson’s algorithm — now RFC 6298 — folds it into a smoothed average and a variance estimate:

RTTVAR = (1 - 1/4) · RTTVAR + 1/4 · | SRTT - R |
SRTT   = (1 - 1/8) · SRTT   + 1/8 · R

RTO    = SRTT + max(G, 4 · RTTVAR)      clamped to [200ms, 120s] on Linux

The variance term is the part people skip, and it is the part that matters. A path with a rock-steady 50 ms RTT and a path that jitters between 10 ms and 90 ms have the same mean. The first deserves an RTO just above 50 ms; the second would retransmit constantly at that value. Multiplying the deviation by four buys headroom proportional to how unpredictable the path actually is.

There is a subtlety in taking those samples. If a segment was retransmitted, and an ACK for it arrives, you cannot tell whether it is acknowledging the original or the copy — and the two interpretations give wildly different RTTs. Karn’s algorithm resolves it by refusing to take an RTT sample from any retransmitted segment at all. The simulator above marks those ACKs, so you can watch SRTT freeze during a recovery.

Three duplicate ACKs mean loss, and waiting for the RTO is a waste

When a segment goes missing, everything after it still arrives. The receiver cannot advance rcv_nxt past the hole, so each arrival produces another ACK carrying the same number. Those duplicate ACKs are information: they prove packets are still flowing, which means the path is alive and one specific segment is missing.

Waiting a full RTO to act on that is pure latency. Fast retransmit takes three duplicate ACKs as sufficient proof and resends immediately. Why three and not one? Because reordering also produces duplicate ACKs, and a network that delivers packets slightly out of order is normal. Three is the empirical threshold where reordering becomes unlikely enough to act on.

Modern Linux does better still. With SACK enabled — and it is, on everything — the receiver reports exactly which ranges it holds, instead of just repeating a high-water mark. The sender then retransmits precisely the missing ranges rather than guessing. RACK-TLP, the default loss detection since 4.18, goes further and uses time rather than dup-ACK counts, which recovers correctly even when reordering is heavy.

Part 3: Flow control — the receiver sets the pace

Reliability tells you what happens when a byte is lost. It says nothing about how fast to send in the first place. There are two separate answers to that question, they come from two different places, and confusing them is the single most common source of wrong TCP intuition.

The first is flow control, and it protects the receiver. When bytes arrive, the kernel puts them in that socket’s receive buffer, where they sit until the application calls read(). If the sender is faster than the application, the buffer fills. There is nowhere else for the data to go, so TCP has to be able to say “stop.”

The mechanism is one field in the header. Every ACK carries a receive window — rwnd — which is the free space left in that buffer. The rule for the sender is absolute: never have more unacknowledged bytes in flight than the last advertised rwnd.

↳ the receiver decides how fast you may send

0 B
buffered
64 KB
rwnd advertised
0 B
unacked in flight
0
zero-window probes

The window field is 16 bits, so it tops out at 65535 bytes. RFC 1323 negotiates a left-shift once, in the SYN — at 80ms RTT this window caps you at 6.6 Mbit/s no matter how fat the link is.

app drain rate3.3 KB/tick

the sender pushes up to 3.7 KB/tick — drag the drain below that and watch the tank win

network → buffer

+0 B/tick

app read() ← buffer

−0 B/tick

delivered total

0 B

0%of 64 KB

last ACK the sender received

ack=… win=65535

true free space right now is 64 KB — the sender is working from a number that is half an RTT stale

sender is link-limited: there is window headroom left over
press play, then starve the application

Zero window, and the timer that stops it deadlocking

Starve the application in that widget and the tank fills, rwnd reaches zero, and the sender stops dead. This is not an error condition — it is flow control working exactly as designed, and it is also why a slow consumer shows up as a stalled producer several services away.

But it sets up a deadlock. The sender is waiting for a window update. The window update is an ACK. ACKs are only sent in response to data. The sender cannot send data. If the one window-update ACK that would have restarted things is dropped, both ends wait forever.

TCP breaks it with the persist timer. While the window is zero, the sender periodically transmits a one-byte probe purely to force the receiver to answer with a fresh ACK, and therefore a fresh window. It is a small, deliberate protocol violation that turns a permanent deadlock into a bounded delay.

Sixteen bits was not enough

The window field in the TCP header is 16 bits: 65535 bytes, maximum. That was generous in 1981 and is now a serious limit, because throughput on a lossless path is bounded by window over round-trip time:

throughput ≤ window / RTT

    64 KB / 1 ms    =   524 Mbit/s     same rack, fine
    64 KB / 40 ms   =    13 Mbit/s     London to Frankfurt
    64 KB / 150 ms  =   3.5 Mbit/s     London to Sydney

  1 MB  / 150 ms  =    56 Mbit/s     the same link, wscale 4

That third line is why a cross-continent transfer on a 10 Gbit link can crawl at 3.5 Mbit/s with nothing wrong anywhere. The pipe is not full; the window is too small to fill it. This quantity — bandwidth × delay — is the bandwidth-delay product, and it is the amount of data that has to be in flight to keep a path busy.

RFC 1323 fixed it with the window scale option: a shift count, negotiated once in the SYN, applied to every window field afterwards. wscale 7 multiplies by 128, taking the ceiling to 8 MB. It has to be in the SYN because both sides must agree before any window is ever interpreted — there is no way to renegotiate later, and a middlebox that strips the option silently pins you to 64 KB for the life of the connection.

Part 4: Congestion control — the sender guesses the network’s limit

Flow control stops you from overwhelming the receiver. Nothing so far stops you from overwhelming the path. The receiver may have a gigabyte of buffer and still be behind a switch whose queue is a few hundred packets deep.

This is the harder problem, because there is no field for it. No router tells you its queue depth. The sender has to infer the network’s capacity from the only signal it gets: whether packets arrive. The variable it maintains for that guess is the congestion window, cwnd, and it is entirely local — it appears in no header anywhere.

So a sender is bounded by both windows at once:

bytes in flight  ≤  min( cwnd, rwnd )
                      ↑     ↑
                      │     └── the receiver's buffer, in every ACK header
                      └──────── the sender's private estimate of the network

When throughput is disappointing, the first useful question is which of those two is binding. They have completely different fixes: rwnd-limited means tune buffers, cwnd-limited means you are losing packets somewhere.

↳ cwnd: how TCP finds the network’s limit

10.0
cwnd (MSS)
64
ssthresh
80
rwnd (MSS)
10
min(cwnd, rwnd)

phase: slow start · round 0 · effective window 10 MSS ≈ 1 Mbit/s at 80ms RTT

0255075100rwndRTT roundssegments in flight (MSS)
Reno cwnd ssthresh
rwnd ceiling80 MSS

cwnd is the sender’s own guess about the network. rwnd is the receiver’s hard limit. You send min() of the two, so pull this down and cwnd stops mattering.

press play — cwnd starts at IW10 and doubles every RTT until it hits ssthresh

Slow start is not slow

A new connection knows nothing about the path, so it starts small and probes upward — but it probes exponentially. Every ACK increases cwnd by one MSS, which means it doubles every round trip. Linux starts at 10 segments (RFC 6928), so a connection reaches roughly 640 segments in flight in six round trips.

The name is historical. It is called slow start because it starts slow, not because it is slow — it is the most aggressive growth phase TCP has, and for short-lived connections it is the only phase that ever runs. A 40 KB HTTP response finishes inside slow start and never touches congestion avoidance at all. This is exactly why connection reuse matters so much for web latency: a warm connection has a large cwnd already, a cold one has ten segments.

AIMD: the shape of the sawtooth

Doubling forever obviously ends badly, so ssthresh marks where TCP switches from probing to creeping. Above it, growth becomes additive increase: one MSS per round trip, not per ACK. On loss it does multiplicative decrease: halve.

Additive-increase / multiplicative-decrease is not arbitrary. It is the rule that makes independent senders converge on a fair share of a shared link without talking to each other. Cautious on the way up, decisive on the way down. That is the sawtooth in the chart, and it is TCP’s entire theory of fairness.

Not all losses are equal

Press both loss buttons in the visualizer and watch the difference, because it is the difference between a hiccup and an outage.

  • Three duplicate ACKs — packets are still arriving, so the path is alive and the ACK clock is intact. Reno halves cwnd, sets ssthresh to match, and continues in congestion avoidance. This is fast recovery, and it costs you half your throughput for a few round trips.
  • An RTO — nothing came back at all. TCP has lost its ACK clock entirely and no longer has any evidence about the path. cwnd collapses to one segment and slow start restarts from scratch. On a 100 ms path that is most of a second before you are back to where you were.

That gap is why fast retransmit exists, why SACK matters, and why a tail-loss event — losing the last packets of a response, where there is no subsequent data to generate duplicate ACKs — used to be so brutally expensive. Tail Loss Probe, now standard, exists specifically to convert those RTOs into fast retransmits.

Why Linux does not ship Reno

Reno’s +1 MSS per RTT is far too timid on a fat, long path. To fill a 10 Gbit link at 100 ms RTT you need about 85,000 segments in flight; recovering from one halving at one segment per RTT takes roughly 42,000 round trips, which is over an hour. In that time you will certainly lose another packet, and the window will never get near the ceiling.

CUBIC, the Linux default since 2.6.19, replaces the linear ramp with a cubic function of the time since the last congestion event:

W(t) = C · (t − K)³ + W_max        C = 0.4,  K = ∛( W_max · β / C ),  β = 0.3

The curve is steep immediately after the drop, flattens as it approaches W_max — the window that caused the loss, and therefore the best available estimate of the path’s capacity — and then steepens again to probe past it. Turn on the CUBIC series in the chart and trigger a loss: it spends its time near the ceiling rather than crawling toward it.

Two more worth knowing by name. BBR abandons loss as a congestion signal entirely and models the path’s bottleneck bandwidth and minimum RTT directly, which makes it much better on lossy paths and much better at avoiding bufferbloat — at the cost of being aggressive toward CUBIC flows sharing a link. ECN lets routers mark packets instead of dropping them, so congestion can be signalled without losing anything. Switching is one sysctl:

shell
# what is available, and what is in use
sysctl net.ipv4.tcp_available_congestion_control
sysctl net.ipv4.tcp_congestion_control

# switch the default (BBR needs the fq or fq_codel qdisc to pace correctly)
sysctl -w net.core.default_qdisc=fq
sysctl -w net.ipv4.tcp_congestion_control=bbr

Part 5: The state machine, and what TIME_WAIT costs

Everything so far is what TCP does with bytes. Underneath it, every connection is a finite state machine, and most of the operational surprises are transitions people have never looked at.

↳ every connection is a state machine

active close — you called close() first

connect() / send SYNrecv SYN,ACK / send ACKclose() / send FINrecv ACK of FINrecv FIN / send ACK2×MSL elapsedCLOSEDLISTENSYN_SENTSYN_RECEIVEDESTABLISHEDFIN_WAIT_1CLOSINGCLOSE_WAITFIN_WAIT_2LAST_ACKTIME_WAITCLOSED

step 0 / 6 · socket()

no connection exists. No kernel state has been allocated.

what TIME_WAIT costs at scale

new conns/sec300/s

300/s × 60s = 18,000 sockets parked in TIME_WAIT against 28,232 ephemeral ports. Still inside the port range — but this is per destination pair, and one hot upstream is all it takes.

The two closes are not symmetric

Step through the active and passive scenarios back to back. The side that calls close() first ends up in TIME_WAIT and holds its four-tuple for a full minute after the connection is functionally over. The side that receives the FIN passes through CLOSE_WAIT and LAST_ACK and is completely finished first, holding nothing.

Whoever closes first pays. That one fact explains a whole category of production behaviour: why your load balancer accumulates TIME_WAIT and your backends do not, why moving the close to the other end of a connection changes which box runs out of ports.

CLOSE_WAIT is always your bug

A socket in CLOSE_WAIT means the peer sent a FIN, the kernel acknowledged it, and it is now waiting for your application to call close(). The kernel cannot do it for you — you might still have data to send. There is no timeout on this state.

So sockets stuck in CLOSE_WAIT are never a network problem and never the peer’s fault. They are a leaked file descriptor: an error path that returns without closing, a connection pool that drops a connection without releasing it, a defer that was never written. If ss -tan state close-wait keeps climbing, go look at your error handling, not your network.

Why TIME_WAIT exists, and why you should not disable it

TIME_WAIT looks like pure waste — the connection is over, why hold the tuple? It does two jobs.

First, it absorbs stragglers. A delayed duplicate from the old connection can still be wandering the network. If the same four-tuple were immediately reused, that segment could be accepted as valid data on the new connection. Holding the tuple for twice the maximum segment lifetime guarantees anything still in flight has expired.

Second, it protects the teardown itself. The final ACK might be lost. If it is, the peer retransmits its FIN — and someone has to be there to answer. A socket that vanished immediately would reply with an RST, and the peer would report a connection error on a connection that closed perfectly.

The arithmetic is what hurts. Linux uses a fixed TCP_TIMEWAIT_LEN of 60 seconds, and the default ephemeral range gives you 28,232 ports per destination pair. At 500 connections per second to one upstream, you hold 30,000 tuples — and connect() starts failing with EADDRNOTAVAIL. Drag the slider in the widget to find your own cliff.

The fix is almost never a sysctl. It is to stop creating a connection per request:

go
// The default http.Client keeps 2 idle connections per host. Under any real
// concurrency that means most requests dial, use, and close — and every one of
// those closes parks a tuple in TIME_WAIT for 60 seconds.
transport := &http.Transport{
    MaxIdleConns:        512,
    MaxIdleConnsPerHost: 128,  // the one that actually matters
    IdleConnTimeout:     90 * time.Second,
}
client := &http.Client{Transport: transport, Timeout: 5 * time.Second}

// And read every response body to completion, or the connection is never
// returned to the pool and the tuning above buys you nothing.
resp, err := client.Get(url)
if err != nil {
    return err
}
defer resp.Body.Close()
io.Copy(io.Discard, resp.Body)

SO_REUSEADDR and SO_REUSEPORT do different things

These get cargo-culted together and they are unrelated.

  • SO_REUSEADDR lets bind() succeed on a local address that still has connections in TIME_WAIT. It is what stops “address already in use” when you restart a server. It does not let two live listeners share a port, and it does not reduce the number of sockets in TIME_WAIT by one.
  • SO_REUSEPORT (Linux 3.9+) lets multiple sockets bind the same address and port simultaneously, with the kernel hashing each incoming connection to one of them. It is how you run N worker processes each with their own listening socket and no accept-queue contention.

And a warning on the sysctls people reach for: net.ipv4.tcp_tw_recycle was removed in Linux 4.12. It broke connections from clients behind NAT badly and silently. tcp_tw_reuse still exists and is safe for outbound connections, because it only reuses a TIME_WAIT tuple when timestamps prove the old incarnation is gone. Neither is a substitute for connection pooling.

Part 6: What the Linux kernel actually runs

Everything up to here is the protocol. This is the implementation — the code your write() lands in, the struct your connection lives in, and the hooks you can attach to when you need to see it. Most engineers have never looked at this layer, and it is where every question about TCP performance is eventually answered.

↳ one sk_buff, seven layers

applicationwrite() / send()sk_buff
socket / TCPtcp_sendmsg()
TCPtcp_transmit_skb()
IPip_queue_xmit()
qdiscdev_queue_xmit()
driverndo_start_xmit()
wirewire

send step 1 / 7 · write() / send()

hands a byte range to the kernel. Returns as soon as the bytes are copied into the socket buffer — not when they are delivered, and not when they are ACKed.

interrupt handling
0
hardware interrupts
0
packets delivered

NAPI: the first packet raises an interrupt, the driver masks the rest and the softirq polls the ring for up to netdev_budget (default 300) packets. Interrupt cost is amortised over the batch.

The sk_buff: one struct, every layer

Every packet in the Linux network stack is an sk_buff. The same allocation travels from the NIC driver to the socket and back, and the reason it is designed the way it is comes down to one requirement: no layer may copy the payload.

It holds a pointer to a data area plus separate pointers to where the transport, network and MAC headers live inside it. Adding a header on the way out is skb_push() — move the data pointer backwards into pre-allocated headroom and write there. Stripping one on the way in is skb_pull() — move the pointer forwards. Seven layers of encapsulation, zero memcpy of payload.

That design is also what makes sendfile() and splice() possible. If the payload never needs touching by the CPU, it never needs to enter userspace at all: the pages go straight from page cache to sk_buff to NIC by DMA.

Where your connection lives: struct tcp_sock

Every TCP connection is a struct tcp_sock, which embeds struct inet_connection_sock, which embeds struct sock. Strip away the several hundred fields and what is left is every variable from the previous five sections:

c
/* include/linux/tcp.h — the fields this post has been describing */
struct tcp_sock {
    u32  snd_una;      /* oldest unacknowledged byte                    */
    u32  snd_nxt;      /* next sequence number to send                  */
    u32  rcv_nxt;      /* next byte expected from the peer              */

    u32  snd_wnd;      /* the window the peer advertised to us (rwnd)   */
    u32  rcv_wnd;      /* the window we are advertising to the peer     */

    u32  snd_cwnd;     /* congestion window, in MSS units               */
    u32  snd_ssthresh; /* slow start threshold                          */

    u32  srtt_us;      /* smoothed RTT, in microseconds << 3            */
    u32  mdev_us;      /* medium deviation                              */
    u32  rttvar_us;    /* smoothed mdev — the RTO variance term         */

    struct sk_buff_head out_of_order_queue;  /* segments past a gap     */
    ...
};

Two details worth carrying. srtt_us is stored left-shifted by three, so a raw read has to be divided by eight to be microseconds — the shift is how the fixed-point smoothing avoids floating point in the kernel. And since 5.19, snd_cwnd is accessed through the tcp_snd_cwnd() accessor rather than directly, so out-of-tree code and old blog posts that touch the field will not compile against a current tree.

You do not need a debugger to read any of this. ss -ti dumps it per socket:

$ ss -ti state established '( dport = :443 )'

Recv-Q Send-Q  Local Address:Port    Peer Address:Port
     0  14600  10.0.1.7:54312        10.0.4.9:443
     cubic wscale:7,7 rto:212 rtt:11.4/2.75 mss:1448 pmtu:1500 rcvmss:1448
     cwnd:24 ssthresh:18 bytes_sent:184320 bytes_acked:169720 bytes_received:8214
     segs_out:132 segs_in:96 data_segs_out:126 send 24.4Mbps lastsnd:4 lastrcv:12
     pacing_rate 29.3Mbps delivery_rate 21.1Mbps busy:284ms retrans:0/3 rcv_space:14480

Read that line by line and the whole post is in it. rto:212 and rtt:11.4/2.75 are SRTT and RTTVAR feeding the formula from Part 2 — 11.4 + 4×2.75 ≈ 22 ms, floored to the 200 ms minimum and rounded. cwnd:24 against ssthresh:18 says this connection is in congestion avoidance, past a loss. retrans:0/3 says three retransmissions have happened over its life and none are outstanding now. Send-Q 14600 is bytes sitting in the send buffer that the kernel has not managed to push out.

NAPI: why your NIC does not interrupt per packet

The obvious receive design is one hardware interrupt per arriving frame. At 10,000 packets per second that is fine. At 10 Gbit line rate with small packets — 14.8 million packets per second — it is a livelock: the CPU spends 100% of its time entering and leaving the interrupt handler and never returns to userspace at all.

NAPI inverts it. The first packet raises an interrupt; the handler immediately masks further interrupts from that queue and schedules a softirq. That softirq then polls the RX ring, draining up to a budget of packets in one pass, and only re-enables interrupts when the ring runs dry. Load goes up, interrupt rate goes down, and the cost per packet falls.

On the way through, napi_gro_receive() does Generic Receive Offload: adjacent segments of the same flow are merged into one large sk_buff before the stack above sees them, so tcp_rcv_established() is called once for 40 KB instead of thirty times for 1448 bytes. Toggle NAPI off in the widget and watch the interrupt count against a 64-packet burst.

Socket buffers autotune, and you should let them

Buffer sizing is where most TCP tuning advice goes wrong. The kernel already does it:

shell
$ sysctl net.ipv4.tcp_rmem net.ipv4.tcp_wmem
net.ipv4.tcp_rmem = 4096   131072  6291456    # min  default  max
net.ipv4.tcp_wmem = 4096   16384   4194304

$ sysctl net.ipv4.tcp_moderate_rcvbuf
net.ipv4.tcp_moderate_rcvbuf = 1                # receive autotuning: on

With autotuning on, the kernel grows the receive buffer as it measures the connection’s bandwidth-delay product and shrinks it under memory pressure. A high-BDP connection gets megabytes; an idle one gets kilobytes.

The trap: setting SO_RCVBUF explicitly turns autotuning off for that socket. You are no longer overriding the default — you are overriding the algorithm, permanently, with a constant you guessed once. Raise tcp_rmem’s maximum instead and let the kernel use the headroom when a connection can actually justify it.

eBPF: watching all of it without touching a packet

The kernel exposes stable tracepoints at every interesting TCP event, which means you can observe connection behaviour in production without tcpdump, without a capture file, and without copying payload anywhere.

shell
# every retransmit, live, with the process responsible
bpftrace -e 'tracepoint:tcp:tcp_retransmit_skb {
    printf("%-16s retransmit %s:%d -> %s:%d\n",
           comm, ntop(args->saddr), args->sport, ntop(args->daddr), args->dport);
}'

# who is sending RSTs, and who is receiving them
bpftrace -e 'tracepoint:tcp:tcp_send_reset    { @sent[comm] = count(); }
             tracepoint:tcp:tcp_receive_reset { @recv[comm] = count(); }'

# a histogram of how long connections actually live
bpftrace -e 'kprobe:tcp_set_state { @[arg1] = count(); }'

This is the layer real observability tooling is built on. tcpretrans, tcplife and tcpconnlat from BCC are each a few dozen lines around exactly these hooks, and every eBPF-based agent that reports per-connection latency is doing the same thing: attaching to tcp_set_state and tcp_rcv_established, keeping a map keyed by socket pointer, and emitting on teardown.

The reason it is worth knowing is cost. A packet capture at 10 Gbit is not something you run on a production box during an incident. A tracepoint that increments a counter in a BPF map is nanoseconds, always on, and gives you the one number you wanted — which sockets are retransmitting, right now, and whose code opened them.

Part 7: The parts that page you

Six sections of theory. This is the on-call version.

Nagle plus delayed ACK: the 40 ms stall

Two reasonable optimisations that combine into a pathology. Nagle’s algorithm refuses to send a second small segment while an earlier one is unacknowledged, to avoid filling the network with 41-byte packets carrying one byte of telnet. Delayed ACK holds an acknowledgement for up to 40 ms hoping to piggyback it on a reply, to avoid sending bare ACKs.

Put them on opposite ends of a request/response connection and they deadlock against each other on every single request.

↳ Nagle × delayed ACK = 40ms of nothing

100ms
message complete at
20ms
best case (one-way delay)
5.0×
latency multiplier
application
3 × write(50B)
sender TCP
send 50B (nothing unacked yet)
Nagle holds 100B — waiting for the ACK
ACK in — release 100B
wire
50B in flight
ACK in flight
100B in flight
receiver
50B received — partial message
delayed ACK timer — 40ms of nothing
ACK sent
full message delivered to app
020406080100120

Nagle is waiting for an ACK. The ACK is waiting for a reply. The reply is waiting for the rest of the message. 40ms of dead air, on every single request.

Rule of thumb: set TCP_NODELAY on anything request/response — RPC, HTTP, database drivers, Redis clients. Leave Nagle on for bulk streaming where you would rather have full segments than low latency.

The signature is unmistakable once you have seen it: a latency histogram with a hard spike at 40 ms that no amount of profiling accounts for, because no code is running during it. The fix is TCP_NODELAY, and it belongs on essentially every request/response socket you own — RPC clients, HTTP clients, database drivers, Redis clients. Go sets it by default on net.TCPConn; most C and Python code does not.

Half-open connections and keepalive

If the peer’s machine loses power, no FIN and no RST is ever sent. Your socket stays ESTABLISHED forever, because an idle TCP connection sends nothing and therefore learns nothing. You find out when you next write and the retransmits eventually time out — minutes later.

TCP keepalive exists for this, and its defaults are useless:

shell
$ sysctl net.ipv4.tcp_keepalive_time net.ipv4.tcp_keepalive_intvl net.ipv4.tcp_keepalive_probes
net.ipv4.tcp_keepalive_time   = 7200    # 2 hours of idle before the first probe
net.ipv4.tcp_keepalive_intvl  = 75      # then a probe every 75s
net.ipv4.tcp_keepalive_probes = 9       # 9 failures before giving up

# 7200 + 9×75 = 2 hours 11 minutes to notice a dead peer

Two hours is not a health check. Either set the per-socket options — TCP_KEEPIDLE, TCP_KEEPINTVL, TCP_KEEPCNT — to something like 30/10/3, or run application-level heartbeats, which have the advantage of also proving the remote process is alive rather than just its kernel. gRPC and most modern RPC frameworks do the latter.

Retransmit backoff: why a dead peer takes 15 minutes

Every RTO doubles the next one. Starting at 200 ms: 0.2, 0.4, 0.8, 1.6 … and Linux gives up after tcp_retries2 attempts, which defaults to 15 and works out to roughly 924 seconds — a little over 15 minutes — before write() finally returns ETIMEDOUT.

Fifteen minutes is far longer than any request timeout you have. If your service is holding a connection to a host that has vanished, the kernel will not tell you for a quarter of an hour. Set TCP_USER_TIMEOUT on the socket — it caps how long unacknowledged data may remain outstanding before the connection is failed, in milliseconds, and it is the single most useful socket option almost nobody sets.

Reading the state of the machine

shell
# the summary — where are all the sockets?
ss -s

# the specific states that indicate specific bugs
ss -tan state time-wait  | wc -l      # closing too much: pool your connections
ss -tan state close-wait | wc -l      # leaking fds: your bug, in your error paths
ss -tan state syn-recv   | wc -l      # SYN queue filling: flood, or backlog too small

# per-socket internals: cwnd, rtt, retrans, pacing
ss -ti state established

# system-wide counters — retransmit rate is the number that matters
nstat -az | grep -E 'TcpRetransSegs|TcpOutSegs|TcpExtTCPLostRetransmit|ListenDrops'

One derived number is worth more than the rest: TcpRetransSegs / TcpOutSegs. Below about 0.1% is a healthy network. Above 1% and your congestion window is spending its life halved, which shows up to everyone else as unexplained tail latency.

The bandwidth-delay product, one more time

The last thing to internalise is that a fast link and a fast transfer are different claims. Throughput is bounded by window over RTT, so a 10 Gbit path with 100 ms of latency needs 125 MB in flight to saturate. If your buffers cap the window at 4 MB, you get 320 Mbit/s — 3% of the link — with zero packet loss and nothing to see in any graph.

Distance is not bandwidth. When a transfer between regions is slow, compute the BDP before you blame the network.

The production checklist

  1. Pool connections. Almost every TIME_WAIT problem is a connection-per-request problem wearing a disguise. Set MaxIdleConnsPerHost, or its equivalent, and drain response bodies.
  2. Set TCP_NODELAY on every request/response socket. If you see a 40 ms spike in a latency histogram, this is it.
  3. Set TCP_USER_TIMEOUT so a vanished peer fails in seconds rather than in tcp_retries2’s fifteen minutes.
  4. Do not set SO_RCVBUF. Raise tcp_rmem’s ceiling and let autotuning use it.
  5. Alarm on retransmit rate, not on interface counters. TcpRetransSegs / TcpOutSegs above 1% is a real signal.
  6. Watch CLOSE_WAIT. It only ever grows because of a bug in your code.
  7. Compute the BDP before concluding a long-distance link is slow.
  8. Never enable tcp_tw_recycle. It is gone since 4.12; if you find it in a runbook, delete the line.

Summary: what people get wrong

  1. “Sequence numbers count packets.” They count bytes. That is why TCP is a stream and why your framing is your problem.
  2. “An ACK confirms one segment.” It is cumulative — a high-water mark. Which is why losing an ACK is usually free and losing a segment is not.
  3. “The window controls the send rate.” There are two windows. rwnd protects the receiver, cwnd protects the network, and you send min() of the two.
  4. “Slow start is slow.” It doubles every RTT. It is the fastest growth TCP has, and for short connections it is the only phase that runs.
  5. “A retransmit is a retransmit.” Fast retransmit halves the window. An RTO collapses it to one segment and restarts slow start. The two differ by orders of magnitude.
  6. “TIME_WAIT is a bug to be disabled.” It is what stops a stale segment corrupting a new connection. Fix the connection churn instead.
  7. “CLOSE_WAIT means the peer misbehaved.” It means your application has not called close().
  8. “Bigger buffers are faster.” Setting SO_RCVBUF disables autotuning and usually makes things worse.

None of this is exotic. It is the behaviour of every HTTP request, every database query and every gRPC stream your service has ever made — which is exactly why the 2 a.m. version of you should not be meeting it for the first time.


If reliable delivery at this layer was interesting, the same problem reappears one layer up with different trade-offs: Kafka beyond the basics covers ISR, the high watermark and what acks=all actually buys you — durability built on top of the delivery guarantee described here. PostgreSQL storage internals takes it to disk, where the WAL is doing for a database what sequence numbers and ACKs do for a byte stream. The anatomy of a lightweight monitoring agent's performance applies the same one-layer-deeper habit to a different target — the tool watching all of these systems, debugged through /proc instead of a wire capture. Go channels: what's actually inside hchan runs the same struct-first approach one layer up the stack, where the wait queues are sudogs instead of sk_buffs and the lock is userspace, not the kernel. For distributed agreement on top of all of it, there are interactive visualizers for Raft consensus and consistent hashing.