Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
1 playbook in Observability & Performance
Structured investigations for the failure modes engineers meet in real systems.
Node.js CPU Profiling: Fixing a Stuck Event Loop in Production
A practical guide to diagnosing Node.js CPU spikes using flame graphs, perf, and async hooks. Includes a real incident where a JSON serialization loop blocked the event loop for 12 seconds.
Prometheus Alert Not Firing: A Systematic Debugging Guide
A structured approach to diagnosing why a Prometheus alert rule isn't triggering, covering expression evaluation, staleness, relabeling, and configuration pitfalls.
Sentry Not Capturing Errors: A Debugging Guide
A direct guide to diagnosing why Sentry fails to capture errors in production, covering SDK misconfigurations, network issues, and rate limits.
Distributed Tracing Spans Not Connecting: A Debugging Guide
Guide to debugging missing parent-child connections in distributed traces, covering sampling mismatches, header propagation failures, and clock skew.
Correlation ID Not Propagating Across Asynchronous Boundaries
Debug why your distributed logging correlation ID disappears across async boundaries, thread pools, or event buses — with concrete commands and fix patterns.
Debugging Long GC Pauses in Java: A Practical Guide
A hands-on guide to diagnosing and resolving excessive Java GC pause times in production JVMs.
Long-form thinking on debugging habits, observability, and the systems around the bug.
Logs, Traces, and the Lies They Tell: Debugging a Stuck Queue in Production
A single stuck queue brought down an entire checkout flow. Here's how we traced the failure across services, what the logs didn't say, and the tools that finally showed the truth.
Reading Distributed Traces to Find Latency: A Field Guide
Tracing tools generate a firehose of data. Here's how to filter the signal from the noise and actually find the root cause of high latency.
Why Your P99 Latency Is Lying to You (and What to Use Instead)
The 99th percentile is the go-to metric for tail latency, but it's easy to fool. Here's a real outage caused by trusting P99, and how to build a more honest monitoring stack.
Reading Flame Graphs: The Non-Obvious Patterns That Reveal Real Bottlenecks
Flame graphs are everywhere in performance profiling, but most engineers only scratch the surface. Here are the patterns that actually tell you where your code is slow.
Observability vs. Monitoring: Why Your On-Call Rotation Still Wakes You Up at 3 AM
Monitoring tells you something is broken. Observability lets you figure out why—without guessing, without SSH, without restarting the pod.