Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
1 playbook in Observability & Performance
Structured investigations for the failure modes engineers meet in real systems.
Node.js CPU Profiling: Fixing a Stuck Event Loop in Production
A practical guide to diagnosing Node.js CPU spikes using flame graphs, perf, and async hooks. Includes a real incident where a JSON serialization loop blocked the event loop for 12 seconds.
Prometheus Alert Not Firing: A Systematic Debugging Guide
A structured approach to diagnosing why a Prometheus alert rule isn't triggering, covering expression evaluation, staleness, relabeling, and configuration pitfalls.
Sentry Not Capturing Errors: A Debugging Guide
A direct guide to diagnosing why Sentry fails to capture errors in production, covering SDK misconfigurations, network issues, and rate limits.
Distributed Tracing Spans Not Connecting: A Debugging Guide
Guide to debugging missing parent-child connections in distributed traces, covering sampling mismatches, header propagation failures, and clock skew.
Correlation ID Not Propagating Across Asynchronous Boundaries
Debug why your distributed logging correlation ID disappears across async boundaries, thread pools, or event buses — with concrete commands and fix patterns.
Debugging Long GC Pauses in Java: A Practical Guide
A hands-on guide to diagnosing and resolving excessive Java GC pause times in production JVMs.
Long-form thinking on debugging habits, observability, and the systems around the bug.
Why Your Production Logs Are Lying to You
Logs tell you what the code reported. They almost never tell you what actually happened. Here is the gap, and how to close it.
Debugging Production Issues Without a Debugger: Approaches That Work
Attaching an interactive debugger in production is usually impossible. Here’s how I gather signal, reproduce issues, and restore service using other techniques.
Defensive Logging: Patterns for Surviving Production Data Rot
Most logging advice stops at 'log more'. Here's how to log defensively—handling nulls, encoding, PII, and context propagation before they rot your observability pipeline.
Tracking Down a 200 MB Leak with Python Memory Profilers
A production API was silently leaking 200 MB of RAM every hour. Here's how memory profilers found the culprit—a forgotten NumPy array reference—and how you can apply the same techniques.
Reading CPU Flame Graphs: What the Hot Colors Actually Tell You
Flame graphs are everywhere, but most engineers read them wrong. Here's how to identify real bottlenecks, avoid common misinterpretations, and turn profile data into actionable fixes.
When Logs Lie: The Gaps Between What You Log and What Actually Happened
Logs are the first thing we reach for during an incident. But they can be misleading, incomplete, or outright wrong. Here's when to trust them and when not to.