Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
1 playbook in Observability & Performance
Structured investigations for the failure modes engineers meet in real systems.
OpenTelemetry Metrics Not Exporting: A Field Guide
A direct, actionable guide for debugging OpenTelemetry metrics that fail to export. Covers collector configuration, SDK timing, and common silent failures.
PyTorch CUDA Out of Memory: Diagnosis and Recovery
A practical guide to diagnosing and fixing CUDA out-of-memory errors in PyTorch, covering memory fragmentation, gradient accumulation, and monitoring with nvidia-smi.
Debugging Elasticsearch Slow Queries: From Shard Contention to Circuit Breakers
A field guide to diagnosing and fixing Elasticsearch slow queries, covering shard contention, hot threads, circuit breakers, and real-world mitigation strategies.
Debugging Cumulative Layout Shift (CLS): From Symptoms to Root Cause
A practical guide to diagnosing and fixing CLS issues in web performance, covering real-world causes, tools, and verification.
LCP Slow: How to Diagnose and Fix Largest Contentful Paint Issues
A practical guide to diagnosing and fixing slow Largest Contentful Paint (LCP) in production web apps.
Long-form thinking on debugging habits, observability, and the systems around the bug.
Adding Observability to a 500K-Line Monolith Without a Rewrite
Adding structured logging, distributed tracing, and metrics to a legacy monolith without a full rewrite. Real code examples and a war story from a 500K-line codebase.
Distributed Tracing: Following a Single Request Across Microservices
Distributed tracing lets you follow a request as it hops across services. I'll show you how trace context propagates, why sampling matters, and how tracing helped us debug a 5-second latency spike in production.
Structured Logging in JSON: Fields, Schemas, and Pitfalls from Production
A practical guide to designing JSON log schemas that are queryable, consistent, and actually useful in production — with field recommendations and a war story.
Sentry Error Monitoring: A Practical Setup Guide from a Production Outage
A step-by-step setup guide for Sentry error monitoring, covering source maps, release tracking, grouping rules, and alerting — with a real story of a production outage that taught us the hard way.
Reading Prometheus Alert Rules: A Practical Reference for Debugging Firing Alerts
Alert rules look straightforward until you're staring at a firing alert at 3 AM. This post covers the structure, common traps, and how to extract actionable intent from any rule.
Profiling SQL Queries: Finding and Fixing the Slow 5%
Profiling SQL queries is more than running EXPLAIN. This guide covers practical techniques to find the slowest queries, interpret execution plans, and fix them — with real examples from PostgreSQL and MySQL.