Senior Software Developer / Bharti Airtel

Vaibhav,
one layer down.

I build and operate distributed systems, and write about what is actually happening underneath them — the log on disk, the window on the wire, the tuple in the page.

signal / 001 live
focusdistributed systems & observability
stackGo · Kafka · Elastic · Linux
statusopen to interesting problems
India
Systems, storage, and the layer under the abstraction.

01 / now

Now

Between posts I am mostly reading source, and shipping small Go libraries that scratch an itch from on-call.

01Reading the Linux networking stack — net/ipv4 mostly, which is what the TCP post turned into.

02Maintaining schemadrift, a Go library for catching JSON schema drift inside a consumer instead of after it.

03Contributing WAL-backed persistence upstream to chromem-go, an embeddable vector database for Go.

04Next field note is fighting between the Go scheduler and eBPF for application developers.

02 / stack

Stack

28 things

What I reach for, grouped by the layer it lives at.

languages

GoPythonC++JavaScriptSQLBash

messaging & data

KafkaPostgreSQLRedisElasticsearchNATS

observability

GrafanaKibanaPrometheusElastic StackeBPF

infrastructure

LinuxDockerKubernetesNginxGit

ai & agents

MCPCrewAIChainlitRAG

frontend

ReactAngularTypeScript

03 / experience

Experience

2 roles
current

Senior Software Developer

Bharti Airtel

  • Architected and built a distributed observability platform from the ground up, scaled to 50,000+ agents across the national network, ingesting tens of terabytes of telemetry a day.
  • Built Go/eBPF agents that capture live process and network topology at the host level, surfaced through a full-stack monitoring UI with interactive graphs and dependency views.
  • Designed and shipped low-latency REST APIs (Go/Gin) with Redis-backed caching over a distributed ClickHouse store, powering real-time dashboards and topology views under high query load.
  • Applied deep CPU and memory optimization on the agent side — cutting syscall and file-descriptor churn, reducing /proc scraping overhead, switching serialization formats to cut GC pressure, and tuning goroutine/channel patterns for lower steady-state resource usage.
  • Designed a fine-grained, resource-level RBAC system alongside SAML-based SSO/identity federation layered on API gateway auth, and led rollout of multi-cluster Kubernetes monitoring across the fleet.
  • Built the alerting and anomaly-detection stack end-to-end — Redis-backed state machines, event-driven delivery over NATS JetStream, and statistical/ML anomaly models running per-host, per-metric at scale.
  • Own 15+ distributed Go microservices behind an API gateway, plus the React frontends engineers rely on to investigate incidents in real time.

Specialist Programmer

Infosys

  • Built microservices in Go for an internal workflow platform, wired to Kafka for real-time event streaming, and added caching layers worth a 30% performance improvement.
  • Designed the concurrent execution model — goroutines, channels and mutexes — for parallel jobs and shared resources in the critical stages of the pipeline, with the classical producer-consumer and reader-writer problems as the correctness bar.
  • Built an in-house agentic store and several multi-agent products on MCP servers and internal knowledge sources, giving other teams a reusable agent framework.
  • Shipped the React and Angular dashboards teams used to watch pipelines and debug live executions.

04 / work

Work

5 projects

Things I built that you can poke at, and one or two you cannot.

go library↗

schemadrift

In-process schema drift detection for JSON message streams. Learns the shape of your messages during a warm-up window, then fires a callback on the first message that breaks it — no schema registry, no sidecar, no broker change.

hard part — Inferring a baseline from live traffic that is strict enough to catch a float64 turning into a string, and loose enough not to page you because one producer omits an optional field.

GoKafkaNATSRedis Streams
open source↗

chromem-go — WAL persistence

Write-ahead-log persistence contributed upstream to chromem-go, an embeddable vector database for Go with ~1k stars. Configurable segment size, background sync workers, a real Close(), plus tests and WAL-vs-no-WAL benchmarks.

hard part — Rotation under concurrency — the old code could hand a writer to one goroutine while another closed it, which surfaced as intermittent "file already closed" errors rather than anything reproducible.

GoWALbenchmarks
python package↗

chainlit-crew-adapter

Runs CrewAI crews inside Chainlit properly: crew, task, agent and tool events render as nested Chainlit steps, and a reusable tool lets an agent stop and ask the user a follow-up question mid-run.

hard part — CrewAI emits a flat callback stream and Chainlit wants a nested step lifecycle. Reconciling the two without losing the nesting was most of the work.

PythonCrewAIChainlit
live↗

Consistent hashing visualizer

Step through key distribution, node joins and departures, and why modulo hashing falls apart the moment the node count changes.

hard part — Making virtual nodes legible without turning the ring into noise.

GoReactSVG
live↗

Raft consensus visualizer

Leader election, log replication and network partitions, played out one step at a time.

hard part — Modelling partitions honestly — the interesting bugs only appear when both halves keep running and both think they are right.

GoReact

05 / play

Play

4 open · 3 soon

Calm playgrounds for systems internals at play.string-wise.com ↗. No score, no timer, no way to lose. A few minutes each, and you leave knowing something real.

Kernel Cosmos

beta

Every planet is a Linux process.

processesschedulermemorypage cacheinterrupts

≈5 min · 16 missionsplay ↗

Packet Garden

beta

Grow a constellation, watch it route.

routingOSPFload balancingcachingconvergence

≈5 min · 9 missionsplay ↗

Packet Drift

beta

Fly one HTTPS request across the internet.

DNSTCP handshakeTLSNATBGP

≈3 min · 3 missionsplay ↗

Orbit Synth

beta

A music box that is secretly a CPU scheduler.

CFSround robinniceSCHED_FIFOstarvation

≈5 min · 5 missionsplay ↗

Goroutine Lanterns

soon

Channels as strings of light: hchan, select and deadlocks you can see.

gochannelsselect

≈5 min

Consensus Choir

soon

Raft, as sound. Leader election and log replication you can hear.

raftconsensus

≈5 min

Data Structure Zen Garden

soon

Rake a B-tree. B-trees, hash rings and heaps, arranged calmly.

b-treeshash ringsheaps

≈5 min

06 / writing

Writing

10 entries

Long-form pieces on how systems actually behave, each with interactive visualizers you can break on purpose.

01 / essay

Beyond 65k: what actually limits TCP connections on Linux↗

The 65,535 limit is a myth: a connection is a 4-tuple, not a port. A Go harness pushed to 100,000 connections on a Raspberry Pi 5 shows the real walls in order: file descriptors, ephemeral ports and the bind-before-connect trap, accept queues, conntrack, kernel and Go memory per connection, and TIME_WAIT on proxies. Six interactive visualizers, measured numbers throughout.

tcplinuxnetworkinggoperformance
22 min read
02 / essay

Your retries took down prod. Not the outage.↗

How a 40-second database failover becomes an hour-long outage: synchronized retry storms, backoff and jitter, amplification across layers, what the network drops first (SYN retransmits, accept queues, conntrack, TIME_WAIT), timeout budgets, metastable failure, and the fixes that work: retry budgets, load shedding, idempotency keys. Seven interactive visualizers.

distributed-systemsreliabilitynetworkinggo
26 min read
03 / essay

We diffed our pipeline against Vector's source. Here's what we found.↗

A Go Kafka-to-Elasticsearch pipeline stuck at half of Vector's throughput, diffed line by line against Vector's source and defaults: a fixed semaphore vs. adaptive request concurrency, batch economics, regex constant factors, redundant JSON round-trips, and payload copies — with five interactive visualizers.

gorustperformanceelasticsearchkafka
22 min read
04 / essay

Go channels: what's actually inside hchan↗

The real runtime struct behind every channel, field by field — the ring buffer, the sudog wait queues, select internals, close semantics — used to explain and fix eleven production bugs: leaked goroutines, deadlocked worker pools, a send-on-closed panic race, a select{default:} loop burning a CPU core, and more.

goconcurrencyinternalsproduction
28 min read
05 / essay

The anatomy of a lightweight monitoring agent's performance↗

A lightweight agent meant to be invisible was quietly burning more CPU and memory than it should. Seven unglamorous costs stacked up: file descriptor churn, /proc scraping overhead, JSON serialization, GC pressure, RSS behavior, and unbounded goroutines — each with an interactive visualizer.

performancegoobservabilityinternals
24 min read
06 / essay

TCP from the inside↗

The handshake, retransmission and RTO, flow control versus congestion control, the state machine and TIME_WAIT, what the Linux kernel actually runs — sk_buff, tcp_sendmsg, NAPI, eBPF tracepoints — and the production gotchas that page you at 3am. Seven interactive visualizers.

tcpnetworkinglinuxinternals
30 min read
07 / essay

Kafka beyond the basics↗

The log on disk, consumer group assignment, eager vs cooperative rebalancing, ISR and the high watermark, log compaction, and exactly-once — seven interactive visualizers for the Kafka internals that decide how production behaves.

kafkadistributed-systemsinternals
22 min read
08 / essay

PostgreSQL storage internals↗

What actually happens when you INSERT, UPDATE, and DELETE rows — MVCC, dead tuples, autovacuum, and the cost-based query planner explained with interactive visualizers.

postgresdatabasesinternals
12 min read
09 / visualizer

Consistent hashing — interactive visualizer↗

Step through how consistent hashing distributes keys across nodes, handles node additions and removals, and why it beats modulo hashing for distributed caches.

distributed-systemsgo
interactive read
10 / visualizer

Raft consensus — interactive visualizer↗

Watch leader election, log replication, and network partitions play out step by step in this visual walk-through of the Raft consensus algorithm.

distributed-systemsgo
interactive read

07 / contact

Contact

Happy to talk about distributed systems, storage internals, or anything in the posts above.

Written by Vaibhav Bhardwaj. Everything here is my own opinion, not my employer’s.