Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
0 playbooks in Debugging Craft
No playbooks found
Nothing matches “”. Try a broader term.
Structured investigations for the failure modes engineers meet in real systems.
Diagnosing Missing Cookies in Nuxt SSR Server Context
Nuxt SSR often fails to read cookies server-side if headers are misrouted or missing. This guide shows where things break and how to reliably retrieve cookies in SSR.
Node.js Inspector Debugger: Diagnosing Protocol Disconnections and Async Breakpoints
A practical guide to debugging Node.js with the inspector protocol—covering disconnections, async stacks, and breakpoint failures.
Java Thread Deadlock Detection with jstack: A Practical Guide
Learn how to detect and fix Java thread deadlocks using jstack. This guide covers real-world patterns, command-line usage, and common pitfalls.
Spring Boot Auto-configuration Not Loading: Debugging Conditional Beans
A hands-on guide to debugging Spring Boot auto-configuration failures caused by missing or mismatched conditions, with concrete commands and real-world war stories.
Spring @Transactional Not Rolling Back: Root Causes & Fixes
A senior engineer's guide to why Spring @Transactional fails to roll back — covering checked exceptions, proxy bypass, and connection commit ordering.
Spring Boot Application Startup Failure: A Practical Debugging Guide
A concise guide to diagnosing Spring Boot startup failures with concrete commands, logs, and root causes.
Long-form thinking on debugging habits, observability, and the systems around the bug.
When to Escalate a Bug: A Decision Framework for Junior and Mid-Level Engineers
A practical framework for junior and mid-level engineers to decide when to escalate a bug, with real-world examples and criteria beyond time spent.
Debugging vs. Firefighting: Why Treating Production Incidents as Debugging Sessions Fails
When production goes down, your brain wants to debug. That instinct costs you hours. Here's why firefighting requires a fundamentally different approach, and the specific process I use to switch modes.
Writing an Incident Runbook That Actually Gets Used in Production
A runbook isn't a document—it's a tool. Here's how to write one that reduces MTTR and doesn't embarrass you during the next PagerDuty alert.
Debugging a production incident with your boss on Slack: staying rational when everything is on fire
A personal account of debugging a cascading Redis failure while the VP of Engineering watched, and the mental models that kept the fix from turning into a rollback.
Print Debugging vs Debugger: When printf Wins and When It Fails
Printf debugging is dismissed as primitive, but in many real-world scenarios it's faster and more reliable than a full debugger. Here's the tradeoff with concrete examples from distributed systems and production incidents.
Writing Postmortems That Actually Improve Reliability
A postmortem that blames someone gets filed and forgotten. A postmortem that traces cause without blame becomes a reliability upgrade. Here's how to write the second kind, with a real example from an 18-hour DNS outage.