I gave the same 29-question order-reading exam to two models.
A cheap one (Haiku 4.5) and one that costs about three times as much (Sonnet 5).
The result
Cheap model 28 of 29 questions clean
Expensive model all 28 executed questions clean
I gave the same 29-question order-reading exam to two models. A cheap one (Haiku 4.5) and one that...
I gave the same 29-question order-reading exam to two models.
A cheap one (Haiku 4.5) and one that costs about three times as much (Sonnet 5).
The result
Cheap model 28 of 29 questions clean
Expensive model all 28 executed questions clean

I built a memory system that lets Claude Haiku (the $1/M-token model) answer questions with 100%...

Eleven hours after the model-comparison post went up, a reader named Vinh Nguyen left a comment that...

Two models, one task swap, and the bill flips by 42%. Which LLM is cheapest for you depends on a number most teams have never…

Here's the scoreboard. Same 50 emails, same prompt, same 4-tier...

The problem A cost efficient AI system sends easy work to a cheap model and only escalates hard work...

I ran the $0.14 model against the $0.44 model expecting a close fight. It lost 4-0. Then I looked at what it does to everything…