Buglyst Blog

Learn to debug under pressure.

Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.

( 02 )Deep debugging guides

Structured investigations for the failure modes engineers meet in real systems.

Browse all guides
Guide

Pulumi State Drift: How Refresh Lies and What to Actually Do

A hard-nosed guide to diagnosing and fixing Pulumi state drift when `pulumi refresh` doesn't behave as expected. Covers real root causes, verification steps, and a war story from production.

Cloud
Guide

Ansible Task Not Idempotent: Changed Every Run

An Ansible task that reports 'changed' on every run is breaking idempotency. This guide covers real-world causes beyond check_mode, from missing diff to transient state in shell commands.

Cloud
Guide

Istio Sidecar Envoy 503: Debugging Upstream Connection Failures

A systematic guide to diagnosing 503 responses from Istio sidecar proxies, covering upstream cluster issues, TLS mismatch, and routing misconfigurations.

Kubernetes
Guide

Linkerd mTLS Handshake Failures: A Practical Debugging Guide

Linkerd mTLS connections failing? This guide covers diagnosing broken TLS handshakes, certificate mismatches, and policy misconfigurations in Kubernetes.

Kubernetes
Guide

Debugging SWC Transform Errors: Syntax Errors in Rust-Based JS/TS Transforms

A practical guide to diagnosing and fixing syntax errors when using SWC for JavaScript/TypeScript transforms. Covers common causes like mismatched decorators, unsupported syntax, and misconfigured .swcrc.

Build tools
Guide

Metro Bundler 'Unable to Resolve Module' — Real Debugging & Fix Strategy

A practical guide to debugging Metro Bundler's 'Unable to resolve module' errors in React Native projects, covering cache, aliases, symlinks, and native module misconfigurations.

Build tools
( 03 )Engineering articles

Long-form thinking on debugging habits, observability, and the systems around the bug.

Browse all articles
Article

How to Read Error Messages: A Debugging Protocol

Most engineers skim error messages and miss the signal. Here's a repeatable protocol for extracting maximum information from stack traces, log lines, and crash dumps — with a real incident walkthrough.

Debugging fundamentals
Article

When to Escalate a Bug: A Decision Framework for Junior and Mid-Level Engineers

A practical framework for junior and mid-level engineers to decide when to escalate a bug, with real-world examples and criteria beyond time spent.

Engineering process
Article

Debugging vs. Firefighting: Why Treating Production Incidents as Debugging Sessions Fails

When production goes down, your brain wants to debug. That instinct costs you hours. Here's why firefighting requires a fundamentally different approach, and the specific process I use to switch modes.

Engineering process
Article

Writing an Incident Runbook That Actually Gets Used in Production

A runbook isn't a document—it's a tool. Here's how to write one that reduces MTTR and doesn't embarrass you during the next PagerDuty alert.

Engineering process
Article

Debugging a production incident with your boss on Slack: staying rational when everything is on fire

A personal account of debugging a cascading Redis failure while the VP of Engineering watched, and the mental models that kept the fix from turning into a rollback.

Debugging mindset
Article

Reading Distributed Traces to Find Latency: A Field Guide

Tracing tools generate a firehose of data. Here's how to filter the signal from the noise and actually find the root cause of high latency.

Observability