Most agent failures I've debugged weren't reasoning failures. The model reasoned fine. It just picked the wrong tool, because the tool description didn't tell it what it needed to know.

This is an under-discussed problem, and it gets worse the more integrations you add. Here's what we've learned.

The setup

Say your agent has four calendar-ish tools available:

[