The question

When you give an AI agent tools to complete a real business task, how much of what it does is the task, and how much is just the agent finding its footing — discovering the schema, pulling raw rows into context, re-reading them, hoping it didn't miss a field?

We built a paired benchmark to measure that gap directly, using Foundgine, an open-source .NET semantic execution layer, against a conventional "give the agent raw application tools" flow. Same task, same data, same required final state. Only the execution boundary changes.

The scenario

A banking customer-review task, deliberately not trivial: