Buglyst Blog
Learn to debug under pressure.
Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.
17 playbooks · 509 guides · 95 articles · 12 linked practice labs ·skip to practice
Fast pattern recognition for the production failures engineers see most often.
0 playbooks in Debugging Craft
No playbooks found
Nothing matches “”. Try a broader term.
Structured investigations for the failure modes engineers meet in real systems.
Debugging java.lang.StackOverflowError from Recursion
A practical guide to diagnosing and fixing StackOverflowError caused by unbounded recursion in Java, including thread stack analysis, JVM flags, and code patterns that actually cause this.
Turborepo Cache Not Hitting – Debugging Remote and Local Cache Misses
A practical guide to diagnosing why Turborepo's cache is not hitting, covering both local and remote cache misses with concrete commands and root causes.
Nx Affected Command Not Detecting Changes: Debugging Guide
When `nx affected` commands fail to see your changes, the cause is almost always a mismatch between the git base and the project graph. This guide shows you exactly how to diagnose and fix it.
Husky Git Hook Not Running: Why Your Pre-commit or Commit-msg Hook is Silently Skipped
Diagnose why Husky hooks are not executing despite being installed. Covers .git/hooks issues, path problems, and permission errors.
React Native iOS Pod Install Error: Xcode Build Failures and CocoaPods Conflicts
A practical guide to diagnosing and fixing pod install failures in React Native iOS projects, covering version mismatches, cache issues, and FFI dependencies.
Expo EAS Build Fails: Debugging Build Errors on EAS
A practical guide to debugging Expo EAS build failures, covering common causes, diagnosis steps, and fixes for EAS Build errors.
Long-form thinking on debugging habits, observability, and the systems around the bug.
When to Escalate a Bug: A Decision Framework for Junior and Mid-Level Engineers
A practical framework for junior and mid-level engineers to decide when to escalate a bug, with real-world examples and criteria beyond time spent.
Debugging vs. Firefighting: Why Treating Production Incidents as Debugging Sessions Fails
When production goes down, your brain wants to debug. That instinct costs you hours. Here's why firefighting requires a fundamentally different approach, and the specific process I use to switch modes.
Writing an Incident Runbook That Actually Gets Used in Production
A runbook isn't a document—it's a tool. Here's how to write one that reduces MTTR and doesn't embarrass you during the next PagerDuty alert.
Debugging a production incident with your boss on Slack: staying rational when everything is on fire
A personal account of debugging a cascading Redis failure while the VP of Engineering watched, and the mental models that kept the fix from turning into a rollback.
Print Debugging vs Debugger: When printf Wins and When It Fails
Printf debugging is dismissed as primitive, but in many real-world scenarios it's faster and more reliable than a full debugger. Here's the tradeoff with concrete examples from distributed systems and production incidents.
Writing Postmortems That Actually Improve Reliability
A postmortem that blames someone gets filed and forgotten. A postmortem that traces cause without blame becomes a reliability upgrade. Here's how to write the second kind, with a real example from an 18-hour DNS outage.