A practical, no-fluff field guide to the failure modes that actually bite teams shipping LLM and agent systems in 2025–2026 — and the concrete techniques that address each one.

Every section follows the same shape: What goes wrong → Why it happens → How to fix it → A quick checklist. Skim the fixes, bookmark the checklists.

Grounded in recent work from Anthropic — Effective context engineering for AI agents & Building effective agents, Cognition/Devin — Don't Build Multi-Agents, Meta AI — Agents Rule of Two, Simon Willison — The Lethal Trifecta & prompt-injection research, Chroma — Context Rot, and Nasr, Carlini, et al. — The Attacker Moves Second — plus the hard-won operational lessons everyone rediscovers the hard way.

Companion reads: 🏗️ Building High-Quality AI Agents — A Comprehensive, Actionable Field Guide 📚 (the how to build counterpart to this guide's what breaks), 🤖 SWE-agent — Deep Dive & Build-Your-Own Guide 📘 (ACI design and tool ergonomics that prevent §10 tool-misuse failures), 🙌 OpenHands — Deep Dive & Build-Your-Own Guide 📚 (the event-sourced kernel and autonomy model behind §8 and §13), 🦊 GoClaw Deep Dive 🤖 — A Builder's Guide to a Multi-Tenant AI Agent Platform 📘 (multi-tenant security and provider resilience for §14–15 and §19), 🔮 Hermes Agent — Deep Dive & Build-Your-Own Guide 📘 (cache-stable prompts, progressive-disclosure memory, and the self-improving loop that addresses §4 and §8), and 🏗️ Building Production-Grade Fullstack Products with AI Coding Agents 🤖 — A Practical Playbook 📘 (end-to-end deployment discipline — evals, PR gates, monitoring — that closes §16 and §17).