TL;DRFrontier AI models are getting better at coding and agentic tasks but worse at general writing. Wizard Labs founder Maz Ahmadi reports measurable regression on client-specific prose benchmarks across model upgrades. With 44% of organizations scaling AI enterprise-wide but only 37% seeing EBIT impact (McKinsey), the gap between deployment and value is widening. His prescription: build custom evaluation suites before development begins, and stop selecting models based on public benchmarks.

Every large language model existing today is getting worse at writing, and almost nobody is measuring it.

The industry keeps asking which model is smartest, but the bigger question is: smartest at what?

As tech giants chase lucrative enterprise contracts, they are optimizing frontier models for coding, logic, and autonomous agent workflows. In the process, general prose has become an afterthought. On our client-specific benchmarks, my team noticed a regression that should alarm anyone using AI for communication. As models upgrade, their writing performance is actively declining.

That may sound like an esoteric problem for engineers and content teams. It is not. AI now sits inside the daily work of millions of people. Nearly nine in ten companies now use AI regularly in one business function: drafting reports, answering customers, preparing legal documents, writing code, analyzing research, and making decisions. If the underlying models are being optimized for a different definition of performance, everyone inherits the consequences.