Buglyst Blog

Learn to debug under pressure.

Playbooks for fast pattern recognition, guides for the full investigation, and articles for the engineering judgment around the edges.

( 02 )Deep debugging guides

Structured investigations for the failure modes engineers meet in real systems.

Browse all guides
Guide

Pulumi State Drift: How Refresh Lies and What to Actually Do

A hard-nosed guide to diagnosing and fixing Pulumi state drift when `pulumi refresh` doesn't behave as expected. Covers real root causes, verification steps, and a war story from production.

Cloud
Guide

Ansible Task Not Idempotent: Changed Every Run

An Ansible task that reports 'changed' on every run is breaking idempotency. This guide covers real-world causes beyond check_mode, from missing diff to transient state in shell commands.

Cloud
Guide

Istio Sidecar Envoy 503: Debugging Upstream Connection Failures

A systematic guide to diagnosing 503 responses from Istio sidecar proxies, covering upstream cluster issues, TLS mismatch, and routing misconfigurations.

Kubernetes
Guide

Linkerd mTLS Handshake Failures: A Practical Debugging Guide

Linkerd mTLS connections failing? This guide covers diagnosing broken TLS handshakes, certificate mismatches, and policy misconfigurations in Kubernetes.

Kubernetes
Guide

Debugging SWC Transform Errors: Syntax Errors in Rust-Based JS/TS Transforms

A practical guide to diagnosing and fixing syntax errors when using SWC for JavaScript/TypeScript transforms. Covers common causes like mismatched decorators, unsupported syntax, and misconfigured .swcrc.

Build tools
Guide

Metro Bundler 'Unable to Resolve Module' — Real Debugging & Fix Strategy

A practical guide to debugging Metro Bundler's 'Unable to resolve module' errors in React Native projects, covering cache, aliases, symlinks, and native module misconfigurations.

Build tools
( 03 )Engineering articles

Long-form thinking on debugging habits, observability, and the systems around the bug.

Browse all articles
Article

Mocking External APIs in Tests: Why stubbing HTTP is harder than it looks

Stubbing HTTP calls seems easy until a subtle mismatch between your mock and the real API causes a production incident. Here's what I learned the hard way.

Testing
Article

How I Diagnosed a 1-in-1000 Race Condition in Our CI Pipeline

A race condition that only appeared once every thousand CI runs took weeks to track down. Here's how we finally caught it, what tools helped, and what we changed to prevent similar bugs.

Testing
Article

How Database Indexes Work: A Debugging Story

A practical guide to understanding database indexes through the lens of debugging a real production outage caused by a missing covering index.

Database
Article

Reading PostgreSQL EXPLAIN ANALYZE Output: A Practical Guide with Real Query Examples

EXPLAIN ANALYZE is the single most important tool for diagnosing slow queries. Here's how to read the output, spot common pitfalls, and fix them with real-world examples.

Database
Article

When caches lie: debugging stale data in distributed systems

Cache invalidation is often cited as one of the two hard problems in CS, but the daily reality is subtler: partial staleness, clock drift, and silent evictions. This post walks through real debugging techniques for stale data in Redis, Memcached, and CDN layers.

Database
Article

Debugging Cache Invalidation Failures: A Case Study with Redis and PostgreSQL

Cache invalidation sounds simple — write-through, TTLs, done. But in practice, silent failures hide in race conditions, connection pools, and stale read replicas. Here's a real debugging story with Redis and PostgreSQL that taught me how to find them.

Database