The pitch for agentic AI on enterprise data is genuinely appealing: instead of a chatbot that answers one question at a time, you get something that plans, queries, checks its own work, calls a second tool, and comes back with a finished piece of analysis — or actually does the thing. Reconcile these two systems. Find out why yesterday's pipeline was late and open a ticket. Pull the accounts at churn risk and draft the outreach list.

The demos work. What determines whether the same thing survives contact with your actual data estate has very little to do with which model you picked, and almost everything to do with what's underneath it. Here's the honest list.

What "agentic" actually changes

A single-shot LLM feature — summarize this, answer that — fails visibly. It returns something wrong, a human reads it, and the human moves on. An agent fails differently: it takes a wrong step early, then builds four more steps on top of it, and the output is a confident, internally consistent artifact that's wrong in a way nobody can see without redoing the work.

Two properties drive most of the difficulty: