The number everyone quotes is the wrong one
Ask a room of engineers how many TCP connections one Linux box can hold, and someone says 65,535. Ports are 16 bits, so that must be the ceiling. It sounds right. It is wrong, and it is wrong in a way that hides the limits you actually hit.
A server with one listening port can hold millions of connections. What stops it is a list of other things: file descriptors, kernel memory, the connection-tracking table, accept queues, and, on the client side, a pool of ephemeral ports that is much smaller than 65,535. Each of these is a separate wall with its own error message, and they show up in a predictable order.
To see them in that order, I built a small harness: a Go server that holds connections, a Go client that opens them from many source IPs, and scripts that sample /proc while it runs. The code is at tcp-1m-connections. Then I ran it on the most modest Linux box I had, a Raspberry Pi 5, to make the walls show up early.
↳ measured on a Raspberry Pi 5, stock sysctls
- one source IP, one destination, default port range28,229 connections, then EADDRNOTAVAIL
- same, after a run left 28k sockets in TIME_WAIT, bind() before connect()EADDRINUSE at 1,446 connections
- 10 source IPs, ramp 10k → 100k at 10,000/s100,000 connections, 0 failures, 0 accept-queue overflows
- server RSS per idle connection3.4 KiB (2.0 KiB stack, 1.0 KiB heap)
- kernel slab per socketabout 4.5 KiB, outside the process RSS
- kernel buffer memory (tcp_mem) at 100k idle0 to 64 KiB in total
- conntrack entries per connection, on loopback1
4 cores, 8 GB, 16 KB pages, Linux 6.18, Go 1.24. Client and server on the same box over loopback.
Every number in that table is from a real run. The Pi tops out well before a million, so where this post talks about 1M, the numbers are projections from the measured per-connection costs, and they are labeled that way. The full 1M run on two servers is next, and the repo is set up for it.
1. A connection is a 4-tuple, not a port
The kernel identifies a TCP connection by four values: (source IP, source port, destination IP, destination port). When a packet arrives, the kernel hashes those four values and looks up the socket in a table called the established hash (ehash). Two connections are different if any of the four values differ.
On the server, every connection has the same destination: the server's IP and its listening port, say :8080. The connections differ only in the client's IP and port. One listening socket can, in principle, tell apart 232 × 216 different peers. The port number on the server side is not a budget. It is a label.
On the client, the picture flips. Toward one server ip:port, from one source IP, three of the four values are fixed. Only the source port varies, and the kernel takes it from ip_local_port_range, which on Linux defaults to 32768 60999: 28,232 ports. That is the real small number, and it belongs to the client.
Each independent part of the tuple multiplies the ceiling. Try it:
↳ client-side ceiling = ports × src IPs × dst IPs × dst ports
The server side has no such product to worry about. Its connections all share one destination (IP, port), and each one is unique through the client's (IP, port). That is up to 2³² × 2¹⁶ different peers per listening socket.
With defaults, one client IP gets you to 28,232 connections against one server port. Getting to a million takes about 36 source IPs, or 16 with a widened range, or fewer IPs and several server ports. None of this has anything to do with 65,535.
2. The harness
The server is the simplest design Go offers, and the one most Go services use: one goroutine per connection, blocked in Read. Its per-connection cost is the number worth measuring, because it is the cost most production code actually pays.
func handle(c net.Conn) {
defer func() {
c.Close()
st.active.Add(-1)
}()
buf := make([]byte, *bufSize) // 256 bytes by default
for {
n, err := c.Read(buf)
if n > 0 && *echo {
if _, werr := c.Write(buf[:n]); werr != nil {
return
}
}
if err != nil {
return
}
}
}The client takes a list of targets, a list or range of source IPs, a connection rate, and a list of stages. For each stage it opens the missing connections once each, at the given rate, and records the first error it sees. It does not retry failed connects, so a stage that falls short tells you exactly where and why.
One trick makes the single-box setup work. On Linux, the whole 127.0.0.0/8 block is local, not just 127.0.0.1. So -src-ips 127.0.0.10-127.0.0.19 gives the client ten source IPs with no setup at all:
./bin/server -addr 127.0.0.1:8080 -csv server-stats.csv &
./bin/client -target 127.0.0.1:8080 \
-src-ips 127.0.0.10-127.0.0.19 \
-stages 10000,25000,50000,100000 -rate 10000Loopback is not a fair test of the network, and both processes share the same CPU and memory. It is a fair test of every limit in this post except NIC throughput, which idle connections do not use.
3. Wall: file descriptors
Every socket is a file descriptor. A process can hold at most RLIMIT_NOFILE of them. There are three numbers stacked on top of each other here, and each one caps the one below it:
- Soft limit, the one that applies. Many distributions default it to 1,024. A C, Python, or Java server that does not raise it gets
accept: too many open files(EMFILE) at about a thousand connections. - Hard limit, the most an unprivileged process can raise its soft limit to. systemd sets it to 524,288 by default.
fs.nr_open, the most anyone can set the hard limit to. The kernel default is 1,048,576. systemd raises it at boot on many distributions; on the Pi it is 1,073,741,816.
There is a fourth, fs.file-max, which caps open files across the whole system. When it runs out, every process on the box gets ENFILE, not just yours. On modern kernels it defaults to a very large number.
Go hides the first number from you. Since Go 1.19, the os package raises the soft limit to the hard limit at startup. The server in the harness prints both, and on the Pi it starts with 524,288. For a Go server, the hard limit is the one that decides when accept fails.
# what the running process actually has
grep 'open files' /proc/$(pgrep -nx server)/limits
# raise it for a systemd service
[Service]
LimitNOFILE=2097152
# fs.nr_open must be at least as high, or setting the limit fails
sysctl -w fs.nr_open=2097152When accept does fail with EMFILE, the worst thing a server can do is retry immediately. The connection is still in the accept queue, so the next accept fails the same way, and the loop burns a CPU core doing nothing. The harness server backs off from 5 ms up to 1 s instead, the same thing net/http does.
On a single box, remember that each connection costs two descriptors: one in the client and one in the server. They are in different processes, so each process needs its own limit, and fs.file-max needs room for both.
4. Wall: ephemeral ports
This is the wall the 65,535 myth is a blurry picture of. It is real, it is smaller than 65,535, and it only applies to the side that calls connect().
With one source IP and one destination, the client in the harness connected 28,229 times and then got:
first connect error at established=27403 attempts=28427:
EADDRNOTAVAIL: dial tcp 127.0.0.2:0->127.0.0.1:18080:
connect: cannot assign requested address28,229 is the default range of 28,232 minus three ports other sockets were using. (The established count in the log line lags a little because a thousand connects were in flight.) The remaining 1,771 attempts all failed, and they took about nine seconds to do it. Each failing connect() scans the whole port range looking for a free port before it gives up, so a client near exhaustion is slow and burns CPU.
The fixes are the multipliers from section 1:
- Widen the range:
net.ipv4.ip_local_port_range = 1024 65535gives 64,512 ports. - Add source IPs to the client.
- Connect to more destination IPs or ports. The server can listen on
:8080,:8081,….
If you widen the range down to 1024, also set net.ipv4.ip_local_reserved_ports for your server's ports. Otherwise a client on the same host can take 8080 as an ephemeral port before the server binds it.
4.1 The trap in "add source IPs"
To use a specific source IP, a client calls bind(ip, port=0) before connect(). Port 0 means "pick one for me". The kernel picks it right there, during bind(), before it knows where the socket will connect to.
So the kernel has to pick a port that is free for every possible destination. A port used by a connection to server A cannot be handed out for server B, even though the two 4-tuples would be different. Each source IP is capped at the size of the port range in total, across all destinations. And bind() also refuses ports still held by sockets in TIME_WAIT.
Linux 4.2 added a socket option for exactly this: IP_BIND_ADDRESS_NO_PORT. With it, bind() records the IP and leaves the port at 0. The kernel picks the port in connect(), when it knows the full 4-tuple, so the same port can be reused toward different destinations. Compare the two with a range shrunk to six ports:
↳ one source IP, three destinations, six ephemeral ports
In bind mode, "connect until it fails" stops at 6 connections in total, with EADDRINUSE. With the option set, it stops at 18: six per destination. Close everything, and in bind mode the TIME_WAIT sockets block the ports for every destination until they expire.
The harness reproduces this for real. With the option turned off (-bind-no-port=false), right after a previous run had left about 28,000 sockets in TIME_WAIT on the same source IP, the client got bind: address already in use at 1,446 connections. Same machine, same source IP, same port range. The client sets the option by default:
const ipBindAddressNoPort = 24 // IP_BIND_ADDRESS_NO_PORT, <linux/in.h>
d := &net.Dialer{
LocalAddr: &net.TCPAddr{IP: srcIP},
Control: func(network, address string, rc syscall.RawConn) error {
var serr error
err := rc.Control(func(fd uintptr) {
serr = syscall.SetsockoptInt(int(fd), syscall.IPPROTO_IP, ipBindAddressNoPort, 1)
})
if err != nil {
return err
}
return serr
},
}Notice that the two failures have different error codes. EADDRNOTAVAIL from connect() means the 4-tuple space is full. EADDRINUSE from bind() means you hit the bind-time cap. And EADDRNOTAVAIL from bind() means something else entirely: the source IP is not on this host. The client checks every source IP before it starts, because that last one looks like port exhaustion in a log and is not.
5. Wall: the accept queue
A listening socket has two queues. The SYN queue holds half-open connections that have not finished the handshake. The accept queue holds connections that are fully established and waiting for the application to call accept(). Its length is min(backlog, net.core.somaxconn). Go passes somaxconn as the backlog, so in Go the sysctl is the limit.
When the accept queue is full, the kernel drops the final ACK of new handshakes. The client retransmits its SYN after 1 second, then 3, then 7. The server's CPU looks idle, because the connections never reached it. The counters that show it are TcpExtListenOverflows and TcpExtListenDrops.
This wall is about rate, not count. A million connections that arrive over ten minutes never fill a 4,096-entry queue. On the Pi, the ramp opened 10,000 connections per second against the default somaxconn of 4,096, and ListenOverflows stayed at zero for the whole run. One goroutine calling accept in a loop kept up easily.
Where it does bite is reconnect storms: a load balancer restart, a deploy, or a network blip that drops every client at once. Then they all come back in the same second. That is the same synchronization problem as a retry storm, with the same fix on the client side: jitter. On the server side, raise somaxconn and tcp_max_syn_backlog, and check ss -lnt: a Recv-Q at or near Send-Q on a listening socket means the queue is full right now.
6. Wall: conntrack
This one surprised me. The Pi has no NAT and no firewall rules for the test port. But nf_conntrack is loaded (Docker and most firewall setups load it), and once it is loaded it tracks every connection that passes through netfilter, including loopback. At 100,000 connections there were 97,000 new conntrack entries: one per connection.
The table is capped by net.netfilter.nf_conntrack_max. On the Pi it is 262,144, a default the kernel sized from RAM. So this box would stop at about 259,000 connections, with plenty of descriptors and memory left, and the only clue would be in dmesg:
nf_conntrack: nf_conntrack: table full, dropping packetNew connections time out instead of failing fast, so from the client this looks like a network problem. The count also plateaus at a suspiciously round number, which is the usual hint. Two fixes:
- Raise
nf_conntrack_max, and the hash table size with it (/sys/module/nf_conntrack/parameters/hashsize). Each entry costs a few hundred bytes of kernel memory. - Or tell netfilter not to track the test port at all, in the
rawtable:
# don't track traffic to or from the server port
iptables -t raw -A PREROUTING -p tcp --dport 8080 -j NOTRACK
iptables -t raw -A OUTPUT -p tcp --sport 8080 -j NOTRACK
# on loopback, the client side goes through OUTPUT too
iptables -t raw -A OUTPUT -p tcp --dport 8080 -j NOTRACKBe careful with NOTRACK on a box with stateful firewall rules. Untracked packets have the state UNTRACKED, not ESTABLISHED, so a rule that only accepts --state ESTABLISHED,RELATED will drop them. And no sysctl on the guest changes a cloud NAT gateway or security group, which keep their own tables with their own limits.
7. Wall: memory
If none of the walls above stop you, memory does. This is the wall where the measured numbers matter most, so here is the whole Pi run, sampled once per second:
↳ measured: 0 → 100,000 connections on a Raspberry Pi 5
Each stage took 1 to 5 seconds at 10,000/s. Zero failed connects, zero accept-queue overflows (somaxconn 4096).
The memory of one idle connection lives in two places, and only one of them shows up in your process.
7.1 User space: 3.4 KiB per connection
The server's RSS went from 5 MiB to 334 MiB for 100,000 connections. The Go runtime's own metrics split that up:
- 2.0 KiB of goroutine stack. 196 MiB of stacks for 100,003 goroutines. A goroutine starts with a 2 KiB stack, and blocking in
Readthrough the network poller never needed more. - 1.0 KiB of heap. The
net.TCPConn, itsnetFDand poll descriptor, the closure, and the 256-byte read buffer. - About 0.35 KiB of other runtime overhead.
The read buffer is the part people get wrong. A 4 KiB buffer per connection, a common default, is almost 4 GiB at a million connections, for memory that is almost always empty. Keep it small, or use a pooled buffer that a connection only takes when it has data.
7.2 Kernel: 4.5 KiB per socket, and almost no buffers
The kernel's own allocations for a socket do not count toward any process's RSS. They live in slab caches: struct tcp_sock (over 2 KiB on its own), the struct file, a dentry and inode for the socket, and the epoll entry Go's poller adds. Slab grew by 870 MiB for 200,000 sockets (both ends of 100,000 connections) and 97,000 conntrack entries. That is about 4.5 KiB per socket.
The other kernel cost is socket buffers, the bytes queued for sending or waiting to be read. Linux allocates these on demand, so an idle socket holds almost none. /proc/net/sockstat reported mem between 0 and 4 pages (0 to 64 KiB) in total for all 200,000 sockets during the holds. The tcp_rmem and tcp_wmem defaults are ceilings for busy sockets, not a per-socket reservation.
That changes the moment connections carry data. net.ipv4.tcp_mem caps buffer memory for all TCP sockets together, and it is measured in pages, not bytes. On the Pi the kernel set its upper limit to 48,090 pages. With 16 KB pages, that is 751 MiB. Copy a tcp_mem line from a guide written for x86 with 4 KB pages, and on a 16 KB-page arm64 box you have allowed four times more memory than the guide meant.
7.3 The budget
Put together, one idle connection costs this server about 7.8 KiB. Change the design, the buffer, the traffic, and the count, and see what fits:
↳ memory per connection, server side
- goroutine stack2.0 KiB
- Go heap: netFD, pollDesc, conn state783 B
- user-space read buffer256 B
- other runtime overhead359 B
- kernel: tcp_sock, file, inode, epoll, conntrack4.5 KiB
- kernel: queued send/receive data0 B
1,000,000 connections fit in 16 GB with 7.5 GiB to spare.
At 1M idle connections the projection is about 7.5 GiB on the server side: 3.2 GiB in the Go process and 4.2 GiB in the kernel. A 16 GB server holds that with room to spare. The goroutine stack is the biggest single line in user space, and an event-loop design (one epoll loop, no goroutine per connection) removes it. That is a large code change for a 2 KiB saving per connection, which is worth it at ten million connections and rarely at one million.
One more Go default matters at this scale. net.Listen and net.Dial turn on TCP keepalive with a 15-second period. With a million idle connections, that is about 66,000 keepalive probes per second, and the same number of replies, doing nothing. The harness server turns keepalive off by default and the client uses 30 seconds. Decide which side is responsible for detecting dead peers, and give only that side keepalive.
8. Which wall do you hit first?
All of these are ceilings on the same number. The one you hit is the lowest, and fixing it just shows you the next one. Start from the stock Pi and fix them one at a time:
↳ six ceilings on one number — the lowest one wins
On the stock Pi, the client's single source IP stops you at 28,232. Add source IPs and conntrack takes over at 262,144. Raise that, and the 524,288 descriptor limit is next, then memory. The tuned preset reaches about a million, limited by the 16 source IPs. Its 16 GB of memory would allow about two million.
The exact numbers depend on your box. The order mostly does not. It is also the order of the error messages you will search for, which is the practical reason to know it.
9. TIME_WAIT and the proxy problem
Everything so far was about holding connections. The other way to run out is to churn them, and this is the one that actually pages people.
The side that calls close() first keeps the 4-tuple in TIME_WAIT for 60 seconds. On Linux that time is fixed in the kernel (tcp_fin_timeout is a different timer, for FIN_WAIT_2). While a tuple is in TIME_WAIT, its ephemeral port cannot be used for the same destination.
Now take a reverse proxy that opens a new connection to one upstream ip:port for every request and closes it afterwards. It has 28,232 ports, and each one is busy for 60 seconds after use. So it can open at most 28,232 / 60 ≈ 470 new connections per second, on any hardware. Above that, it runs out of ports while CPU and memory are idle:
↳ proxy → one upstream ip:port, client side closes first
Opening 2k connections/s against a ceiling of 471/s. The box is idle and the proxy is out of ports.
Raising the port range or adding source IPs moves the line, but slowly: 8 source IPs with the full range still only reach about 8,600 new connections per second. tcp_tw_reuse=1 helps more. It lets the kernel reuse a TIME_WAIT port for a new outgoing connection once a second has passed, using TCP timestamps to keep old and new segments apart. Recent kernels already enable it for loopback traffic by default.
The real fix is the "requests per connection" control. With a connection pool that does 100 requests per connection, the same 2,000 requests per second need 20 new connections per second, and TIME_WAIT never matters. In Go that means draining and closing every response body and setting MaxIdleConnsPerHost to your real concurrency. TCP from the inside covers why the closing side holds TIME_WAIT, and the retries post covers how a retry storm turns this into an outage.
One thing not to do: tcp_tw_recycle. It broke clients behind NAT and was removed from the kernel in 4.12. Old tuning guides still recommend it.
10. Watching it happen
At a million sockets, your usual tools become part of the problem. ss -tan | wc -l and netstat walk every socket, which takes seconds and real CPU. The metrics script in the repo reads only files in /proc, which the kernel keeps as running counters:
cat /proc/net/sockstat # TCP: inuse, orphan, tw, alloc, mem (pages)
cat /proc/sys/fs/file-nr # allocated file handles, system-wide
cat /proc/sys/net/netfilter/nf_conntrack_count
nstat -az TcpExtListenOverflows TcpExtListenDrops TcpExtTCPSynRetrans
ls -U /proc/$(pgrep -nx server)/fd | wc -lFor a live view of what the kernel is doing, the repo also has three bpftrace scripts: accepts per second from inet_csk_accept, connect errors by errno from tcp_v4_connect, and state transitions from the sock:inet_sock_set_state tracepoint. The connect script is the useful one for port exhaustion. tcp_v4_connect is where the kernel picks the ephemeral port, so running out shows up there as -99 (EADDRNOTAVAIL):
kretprobe:tcp_v4_connect
/retval != 0/
{
if (-retval == 99) { @errors["EADDRNOTAVAIL"] = count(); }
else { @errors_other[-retval] = count(); }
}It only catches the connect-time failure. A plain bind() before connect() fails earlier, in bind(), and never reaches this function. That is section 4.1 again, from the kernel's side.
11. Run it yourself
The repo has the server, the client, a sysctl profile that snapshots your current values before changing them, a revert script, and a ramp script that collects everything into one directory per run:
git clone https://github.com/VA-ibh-AV/tcp-1m-connections.git && cd tcp-1m-connections
(cd server && go build -o ../bin/server .)
(cd client && go build -o ../bin/client .)
# server host
sudo ./scripts/apply-sysctl.sh tuning/sysctl-1m.conf
./bin/server -addr :8080 -csv server-stats.csv
# client host
sudo ./scripts/setup-ips.sh eth0 10.0.0.11 10.0.0.50 24
./scripts/run-ramp.sh --target <server-ip>:8080 --src-ips 10.0.0.11-10.0.0.50 \
--stages 10000,100000,500000,1000000 --hold 120s --rate 5000
# afterwards, on both
sudo ./scripts/revert-sysctl.shRun it on machines you own. A million connections from a few IPs looks exactly like a SYN flood to anything in between.
The takeaway
65,535 is the number of ports, not the number of connections. A server's connection count is limited by file descriptors, conntrack, and memory, and on a stock box the first two are configuration, not hardware. A client's connection count is limited by the 4-tuple space toward each destination, which is ports × source IPs × destination addresses, and by how fast TIME_WAIT gives ports back.
A Raspberry Pi held 100,000 idle connections at about 7.8 KiB each, with no tuning at all. The walls are real, but each one has a name, an error message, and a counter in /proc. Find out which one is lowest on your box before production does it for you.
Related: TCP from the inside covers the handshake, the state machine, and the kernel path these connections take. Your retries took down prod covers what happens when all these clients reconnect at once. The anatomy of a lightweight monitoring agent covers the cost of reading /proc in a loop.