Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
1 playbook in Observability & Performance
Structured investigations for the failure modes engineers meet in real systems.
Python Memory Leak Profiling: Diagnosing Unbounded Growth in Production
A practical guide to identifying and fixing Python memory leaks using tracemalloc, objgraph, and heapy, with production-tested strategies.
Python logging not showing output: Why your log messages vanish
A practical guide to diagnosing silent loggers in Python, covering root logger, handler levels, propagation, and common misconfigurations that cause output to disappear.
Go pprof CPU & Memory Profiling: Real-World Debugging Tactics
A hands-on debugging guide for Go pprof CPU and memory profiling with real commands, root causes, and a war story.
Diagnosing and Fixing Java OutOfMemoryError: Java Heap Space
A practical guide to diagnosing Java heap OutOfMemoryError in production, covering non-obvious causes like GC overhead, native memory leaks, and permgen/metaspace issues.
Custom Metrics Not Appearing in AWS CloudWatch: A Debugging Guide
Diagnose why your custom CloudWatch metrics are not showing up. Covers common causes like namespace mismatch, timestamp issues, and IAM permissions.
HTTP Connection Timeout: Debugging Network Stalls and Misconfigured Timeouts
A practical guide to diagnosing HTTP connection timeouts caused by network latency, firewall drops, TCP backlog exhaustion, and misconfigured timeout stack.
Long-form thinking on debugging habits, observability, and the systems around the bug.
Logs, Traces, and the Lies They Tell: Debugging a Stuck Queue in Production
A single stuck queue brought down an entire checkout flow. Here's how we traced the failure across services, what the logs didn't say, and the tools that finally showed the truth.
Reading Distributed Traces to Find Latency: A Field Guide
Tracing tools generate a firehose of data. Here's how to filter the signal from the noise and actually find the root cause of high latency.
Why Your P99 Latency Is Lying to You (and What to Use Instead)
The 99th percentile is the go-to metric for tail latency, but it's easy to fool. Here's a real outage caused by trusting P99, and how to build a more honest monitoring stack.
Reading Flame Graphs: The Non-Obvious Patterns That Reveal Real Bottlenecks
Flame graphs are everywhere in performance profiling, but most engineers only scratch the surface. Here are the patterns that actually tell you where your code is slow.
Observability vs. Monitoring: Why Your On-Call Rotation Still Wakes You Up at 3 AM
Monitoring tells you something is broken. Observability lets you figure out why—without guessing, without SSH, without restarting the pod.