Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
0 playbooks in Debugging Craft
No playbooks found
Nothing matches “”. Try a broader term.
Structured investigations for the failure modes engineers meet in real systems.
React Query Not Refreshing Stale Data Despite Invalidations
Diagnose and resolve cases where React Query fails to update stale data, even when invalidations or refetches are expected to trigger.
“Module Not Supported” Errors in Next.js Edge Runtime
A real-world guide to diagnosing and fixing 'module not supported' errors when deploying Next.js projects to the Edge Runtime.
React Portal Allows Click Events to Bubble Outside Intended DOM
When using React portals, click or key events can bubble up to unexpected ancestors, triggering handlers outside their intended boundaries.
Vue Reactive Property Not Updating in Template
Diagnose and resolve cases where Vue's reactive data fails to reflect changes in the UI, including edge cases beyond the classic pitfalls.
Diagnosing Nuxt Hydration Mismatch Errors in Production
Hydration mismatch errors in Nuxt apps break SSR interactivity and can be tough to pinpoint. This guide walks through investigating, reproducing, and fixing these.
Why Your Vue Composition API ref Isn’t Reactive
Vue’s Composition API refs sometimes fail to trigger reactivity. This guide covers subtle causes, practical debugging, and real-world fixes for non-reactive refs.
Long-form thinking on debugging habits, observability, and the systems around the bug.
When to Escalate a Bug: A Decision Framework for Junior and Mid-Level Engineers
A practical framework for junior and mid-level engineers to decide when to escalate a bug, with real-world examples and criteria beyond time spent.
Debugging vs. Firefighting: Why Treating Production Incidents as Debugging Sessions Fails
When production goes down, your brain wants to debug. That instinct costs you hours. Here's why firefighting requires a fundamentally different approach, and the specific process I use to switch modes.
Writing an Incident Runbook That Actually Gets Used in Production
A runbook isn't a document—it's a tool. Here's how to write one that reduces MTTR and doesn't embarrass you during the next PagerDuty alert.
Debugging a production incident with your boss on Slack: staying rational when everything is on fire
A personal account of debugging a cascading Redis failure while the VP of Engineering watched, and the mental models that kept the fix from turning into a rollback.
Print Debugging vs Debugger: When printf Wins and When It Fails
Printf debugging is dismissed as primitive, but in many real-world scenarios it's faster and more reliable than a full debugger. Here's the tradeoff with concrete examples from distributed systems and production incidents.
Writing Postmortems That Actually Improve Reliability
A postmortem that blames someone gets filed and forgotten. A postmortem that traces cause without blame becomes a reliability upgrade. Here's how to write the second kind, with a real example from an 18-hour DNS outage.