The price of processing AI tokens, the basic unit of work for large language models, has fallen to roughly $1.16 to $1.18 per million tokens in early August 2026. That’s the lowest average inference cost recorded this year, and it represents a 43% decline from $2.04 at the end of May.

To put that trajectory in perspective: frontier AI intelligence is now priced at approximately 12% of its March 2023 levels, according to BenchLM indices. Some benchmarks show comparable AI capabilities costing over 280 times less than they did in early 2023.

What’s driving the freefall

Two forces are crushing inference prices simultaneously, and neither shows signs of slowing down.

The first is OpenAI’s decision to slash pricing on its GPT-5.6 Luna models by 80% in late July 2026. After the cuts, input tokens cost $0.20 per million and output tokens run $1.20 per million.