Things that broke in production (and what I learned)
A collection of "oh no" moments from running real services. No sugarcoating.
Running services that process ~1000 GPS frames per second means things break. Here are some of my favorites.
The Kafka rebalance storm
We had a consumer that took too long to process a batch. Kafka thought it was dead, triggered a rebalance, reassigned partitions. The consumer came back, got new partitions, started processing, took too long again. Rebalance. Loop.
The fix was embarrassingly simple: increase max.poll.interval.ms and process smaller batches. Took 3 hours to figure out, 2 minutes to fix.
Lesson: Read the Kafka consumer config docs. All of them. Before you need them at 2am.
The memory leak that wasn’t
Memory kept growing on one service. Classic leak, right? Spent two days adding heap snapshots, profiling, suspecting every dependency.
Turned out it was just Redis connection objects piling up because we weren’t closing them after pub/sub unsubscribe. The connection pool was “working correctly” by design, just holding references we didn’t need.
Lesson: Not everything that looks like a memory leak is a memory leak. Sometimes you’re just holding references you forgot about.
The silent failure
A geofencing consumer stopped processing events. No errors in logs. No alerts. CPU at 5%. Everything looked healthy.
It had gotten stuck on a malformed message it couldn’t parse, and our error handling was catching the exception but not advancing the offset. So it kept retrying the same broken message forever. Silently.
Lesson: Always have a dead letter queue. Always advance the offset on unrecoverable errors. And alert on consumer lag, not just errors.
The timezone bug
Date stored in UTC. Frontend displayed in local time. Report generation used… neither? Some library default that was apparently US Eastern.
Clients in Poland seeing delivery timestamps off by 7 hours. Nobody noticed for two weeks because “eh, timestamps are always weird.”
Lesson: Pick UTC everywhere. Convert only at display time. Never trust a library’s default timezone.
What I actually learned
Every production incident taught me the same thing: the system is only as good as its monitoring. If I can’t see what’s happening, I can’t fix it. Grafana dashboards, consumer lag alerts, structured logging. Not glamorous, but it’s what lets you sleep at night.