The thing that watches everything else was showing up on its own dashboard
A lightweight agent runs on every host in the fleet. Its whole job is to be invisible — scrape a few metrics, ship them somewhere, use as little CPU and memory as the host it's watching can spare. Someone notices the agent's own steady-state CPU baseline is higher than it should be. Not a spike. Not a leak with a clean upward slope you can point at. Just a floor that sits a few points higher than the same agent on a similar fleet a year ago.
First hypothesis: something got slower in the last release — a new metric, a heavier query, a serialization change nobody benchmarked. Tempting, because it's a single thing to find and revert. It's also wrong. Profiling the hot path shows nothing dramatic — no function eating 40% of samples, no obvious regression. The CPU is going somewhere, just not anywhere a flame graph makes look interesting.
The diagnostic path that actually worked didn't start with a profiler. It started with /proc — descriptor counts, syscall rates, the kernel's own view of what the process is doing per second. Each answer opened a question one layer further down: descriptor counts led to syscall rates, syscall rates led to what those syscalls were parsing, parsing led to allocation rate, allocation rate led to GC frequency, GC frequency led to what the memory graphs actually meant, and memory graphs led to how many goroutines were alive at any given moment. Seven costs. None of them individually explained the gap. Together, they were the whole gap.
This is the same habit of going one layer deeper that showed up tracing bytes through the kernel in TCP from the inside — except this time the "wire" is /proc, and the thing you're debugging is your own observability tooling.
Part 1: file descriptor churn — the graph that lies by omission
A scraper reads a handful of files every tick: /proc/stat for global CPU, /proc/[pid]/stat for per-process numbers, /proc/net/dev for interface counters, maybe a couple of log files it tails. The naive version opens each file, reads it, closes it — every single tick. The efficient version opens once and reuses the descriptor. Both look identical on an FD-count graph, because in neither case does the descriptor count actually grow.
The cost isn't a rising number anywhere obvious. It's the syscall rate. Every open() and every close() is a context switch into the kernel — cheap in isolation, expensive at scrape-interval frequency across an entire fleet. lsof -p PID and ls -la /proc/[pid]/fd | wc -l both report a perfectly steady descriptor count. Neither tool sees the churn.
# descriptor count alone won't show churn — it looks the same either way
ls -la /proc/$PID/fd | wc -l
# this is the tool that actually reveals it
strace -c -p $PID
# % time seconds usecs/call calls syscall
# ------ ----------- ----------- --------- ----------------
# 41.2% 0.008812 4 2160 open
# 33.7% 0.007211 3 2160 close
# 19.8% 0.004230 2 2160 read
# 5.3% 0.001134 1 180 epoll_waitPlay with the visualizer below. Watch the FD-count graph stay perfectly flat for both pooled and churn — then look at the syscalls-this-tick number. That's the whole point: FD count alone can't distinguish churn from a healthy steady state. Only the syscall rate can. And a genuine leak looks different from both — the descriptor count itself climbs, because nothing ever calls close() at all.
↳ FD count vs syscall rate — they tell different stories
open file descriptors over time
press start to begin
FD count graph looks IDENTICAL to pooled — steady at 12. But every scrape does 12 open() + 12 read() + 12 close() calls. lsof shows nothing wrong. strace -c -p PID shows the real cost.
The fix was the boring one: hold descriptors open across ticks, reuse connections instead of dialing fresh ones, and reserve strace -c for confirming it actually worked — the syscall counts dropped by roughly the ratio of scrape interval to file count, exactly as the math predicted.
Part 2: /proc scraping overhead — freshness has a CPU price
Even with descriptors pooled, every scrape still does small reads and string parsing — /proc/stat's space-separated CPU jiffies, /proc/net/dev's fixed-width columns, a handful of files per tick, multiplied by however many hosts run the agent. None of it is expensive once. All of it is expensive at a one-second interval, all day, on every host.
The tradeoff is explicit: a shorter interval means fresher data and more syscalls per second; a longer interval means less CPU and staler data by up to half the interval, on average. There usually isn't one right answer for the whole agent — a metric feeding an alert that pages someone needs to be fresh; a static host label read once and cached needs to be read once, ever, not every tick.
Move the slider below. The curve is a plain 1/interval relationship, but it's worth seeing where it flattens: below roughly 300ms the syscall floor alone becomes a measurable CPU cost, and above a few seconds the freshness loss starts to matter more than the CPU it's saving.
↳ polling interval — freshness vs syscall cost
syscalls/sec across the whole interval range
12 metric sources read per tick, 6µs assumed per small read + parse. Below ~300ms the syscall floor alone becomes measurable CPU; above ~5s freshness starts to matter for anything alerting on the data. The right answer is rarely one interval for everything — batch the reads that share a tick, and give cheap, low-priority metrics (disk labels, static host tags) a much slower interval than CPU or memory.
The practical mitigations, in order of how much they helped: batch reads that share a tick into one pass instead of one syscall round-trip per metric, cache anything that changes slower than the interval (hostname, kernel version, mount points), and give each metric its own interval instead of one global tick for everything.
Part 3: serialization — where JSON actually costs you
The agent ships metrics somewhere, and the wire format was JSON — readable, ubiquitous, easy to curl | jq when something looks wrong. None of that is free. JSON encoding in most languages goes through reflection to walk struct fields, allocates strings for every key and every escaped character, and re-parses numbers as text on the way back in. A "small payload" stops being small once you multiply its encode/decode cost by scrape frequency times fleet size.
MessagePack is the same conceptual model — maps, arrays, integers, strings — with no text parsing and no escaping. Same data, binary wire format, smaller and cheaper to produce. There are three separate wins here, worth measuring separately rather than as one blended number: CPU time to encode/decode, bytes on the wire, and — the one that matters most for Part 4 — how many intermediate allocations each encoding produces.
func BenchmarkEncodeJSON(b *testing.B) {
m := sampleMetric()
for i := 0; i < b.N; i++ {
_, _ = json.Marshal(m)
}
}
func BenchmarkEncodeMsgPack(b *testing.B) {
m := sampleMetric()
for i := 0; i < b.N; i++ {
_, _ = msgpack.Marshal(m)
}
}
// go test -bench=Encode -benchmem
// BenchmarkEncodeJSON-8 421339 2840 ns/op 512 B/op 9 allocs/op
// BenchmarkEncodeMsgPack-8 2938451 410 ns/op 64 B/op 1 allocs/opToggle between a flat payload and a nested one below — the gap widens as the struct grows, because reflection cost and allocation count both scale with field count while the binary encoder's cost scales with byte count.
↳ JSON vs MessagePack — same struct, two encodings
fields in this payload
Three separate wins, worth measuring separately: less CPU to encode/decode, fewer bytes on the wire, and fewer allocations — which is the number that actually drives GC frequency (Section 4). The tradeoff: MessagePack payloads aren't human-readable. You lose curl | jq debugging and pay more attention to schema evolution, since there's no field name in the wire format to fall back on.
The honest tradeoff: MessagePack payloads aren't human-readable, so debugging drops to a decoder tool instead of curl | jq, and schema evolution needs more care since there's no field name embedded in the wire format to fall back on if a decoder gets out of sync with an encoder. Worth it for a payload sent thousands of times a second across a fleet; probably not worth it for a config file a human edits by hand.
Part 4: GC pressure — the agent measures its own noise
GC pauses matter more for an agent than for a typical request/response service, for an uncomfortable reason: the agent is often the thing measuring the very latency and CPU numbers its own pauses pollute. A GC pause that stalls the scraper for a few hundred microseconds shows up as a gap or a spike in exactly the data meant to catch gaps and spikes elsewhere.
The misconception worth killing early: heap size is not what drives GC frequency. Allocation rate is. Go's default GC target (GOGC=100) triggers a collection when the live heap has doubled since the last one — a process with a small, completely stable live heap can still GC constantly if it allocates and discards garbage fast enough to keep hitting that doubling point.
GODEBUG=gctrace=1 ./agent
# gc 142 @6.011s 2%: 0.019+1.8+0.006 ms clock, 0.15+0.41/1.6/3.2+0.048 ms cpu,
# 24->26->13 MB, 26 MB goal, 8 P
#
# gc 142 — the 142nd GC cycle since start
# @6.011s — 6 seconds since process start
# 2% — cumulative % of CPU time spent in GC so far
# 0.019+1.8+0.006 ms — stop-the-world sweep termination + concurrent mark + STW mark termination
# 24->26->13 MB — heap size before GC -> at GC finish -> live set after GC
# 26 MB goal — the doubling target that triggered this cycle
# 8 P — GOMAXPROCSruntime.ReadMemStats gives the same story in code — NumGC, PauseTotalNs, Mallocs/Frees — and a pprof heap profile separates two different questions that are easy to conflate: alloc_objects (how many things are being allocated — usually the GC-frequency question) versus alloc_space (how many bytes — usually the "why did RSS jump" question).
Toggle between the before and after allocation rates from Part 3 below — same base live heap in both cases, only the allocation rate changes, and GC frequency moves with it.
↳ allocation rate drives GC frequency — not heap size
GODEBUG=gctrace=1 — heap in use, GC events as dashed lines
press start to begin
GOGC=100 means: trigger a GC when the live heap has doubled since the last one. The live heap here never grows — it's the same 24MB of real state — but a higher allocation rate reaches that doubling point faster, so GC fires more often. Heap size was never the driver. Allocation rate was.
This is the payoff that ties Part 3 to Part 4: switching from JSON to MessagePack wasn't just a CPU and payload-size win. It cut the allocation rate, and cutting the allocation rate cut GC frequency directly. One fix, three metrics moved together.
Part 5: memory and heap behavior — sawtooth vs staircase
Two numbers that look like they should agree often don't: VmRSS from /proc/[pid]/status, and the heap-in-use number pprof reports. RSS can sit well above what pprof says is live, and that's not automatically a bug — Go's runtime doesn't return freed pages to the OS immediately, so RSS tends to track the historical peak rather than the current heap.
The pattern that matters is the shape, not the gap. A heap that oscillates between the same floor and ceiling every GC cycle — a sawtooth — means the live set isn't growing; RSS holding flat slightly above the ceiling is exactly what's expected. A heap whose post-GC floor itself keeps rising every cycle — a staircase — means something is retaining references the live set doesn't need: an unbounded cache, a slice that only ever appends, goroutines that never exit and never release what they're holding.
↳ RSS vs heap-in-use — sawtooth vs staircase
press start to begin
Heap sawtooths between the same floor and ceiling every GC cycle — the live set isn't growing. RSS holds flat slightly above the ceiling, because Go doesn't return freed pages to the OS on every collection. A stable gap here is normal.
The single best tool for turning "the staircase is real" into "here's what's causing it" is a pprof heap diff between two points in time:
curl -s localhost:6060/debug/pprof/heap > heap.0.pprof
sleep 300
curl -s localhost:6060/debug/pprof/heap > heap.1.pprof
go tool pprof -top -base heap.0.pprof heap.1.pprof
# Shows only what grew between the two snapshots — not the whole heap,
# just the delta. This is almost always faster than reading the whole
# profile and guessing.Part 6: goroutines vs buffers — concurrency design, not a speed fix
Goroutine count is a leading indicator worth watching on its own: unbounded growth — one goroutine per connection, one per file watched, one per incoming request with no cap — shows up as scheduler overhead and memory growth well before it shows up as a crash. runtime.NumGoroutine() tracked over time catches this long before an OOM does.
A buffered channel is backpressure, not a speed fix. An oversized buffer doesn't solve a slow-consumer problem — it hides it, turning a throughput mismatch into silent memory growth instead of a visible signal. Three designs, same load:
- Unbounded buffer — memory grows silently, nothing fails until it's too late to fail gracefully.
- Bounded buffer with an explicit drop or block policy — backpressure becomes visible: a full queue, a dropped-item counter, something you can alert on.
- Worker pool — concurrency itself is bounded; excess work queues (bounded) or blocks, but the goroutine count has a ceiling regardless of load.
// unbounded: one goroutine per connection, no ceiling
for conn := range incoming {
go handle(conn) // under sustained overload this count never stops climbing
}
// worker pool: concurrency capped, backpressure explicit
jobs := make(chan Conn, queueCap) // bounded — a full channel is a decision point, not silent growth
for i := 0; i < poolSize; i++ {
go func() {
for conn := range jobs {
handle(conn)
}
}()
}
for conn := range incoming {
select {
case jobs <- conn:
default:
droppedCounter.Inc() // explicit, visible, alertable — not a memory leak in disguise
}
}Play with the incoming-load slider below and switch designs. pprof's goroutine profile shows where goroutines are stuck — blocked on a channel send, a mutex, or I/O — and GODEBUG=schedtrace=1000 gives scheduler-level visibility (runnable goroutines per P, how often the scheduler is stealing work) when the goroutine count alone doesn't explain the CPU cost.
↳ unbounded goroutines vs a bounded worker pool
runtime.NumGoroutine() over time
press start to begin
Every connection spawns its own goroutine. If the consumer can't keep pace with 40/tick incoming, the count only ever grows — this shows up in pprof's goroutine profile as thousands blocked on the same channel send long before anything OOMs.
Part 7: putting it together
None of these six costs was individually dramatic. FD churn cost a few thousand extra syscalls a second. The scrape interval cost a fraction of a percent of CPU. JSON cost some allocations. Those allocations cost some extra GC cycles. The GC cycles and an unbounded worker design cost some extra memory. Stacked on every host in the fleet, all day, that's the floor that came back down.
↳ every fix, one dashboard — the floor came back down
This is the incident state — every cost from the last six sections stacked on the same fleet of hosts.
What stayed in place afterward, specifically so this doesn't quietly reappear: strace -c in the release checklist for anything touching the scrape loop, an alert on allocation rate (not just heap size) so a regression in Part 3 gets caught before it becomes a Part 4 problem, and a goroutine-count alert with a threshold set to the worker pool's ceiling — any sustained excess past that number means the bounded design stopped being bounded somewhere.
The agent went back to being invisible. That was always the actual spec.
Related: one layer further down the stack, TCP from the inside walks the same kind of investigation through sk_buffs and the kernel's send/receive path — different failure domain, same habit of going one layer deeper each time the first answer doesn't fully explain the symptom. Go channels: what's actually inside hchan takes Part 6's unbounded-goroutines point and goes all the way into the channel internals behind it — the same leaked-goroutine shape shows up there as a full incident, traced back to hchan's wait queues field by field. Also: Kafka beyond the basics and PostgreSQL storage internals for what the systems the agent is watching are doing on their own side of the wire.