OpenAI wants its cleverest model to also be its quickest. The company has previewed Ultrafast, a new tier of its API that runs the flagship GPT-5.6 Sol at up to 14 times the usual speed, reaching around 750 output tokens a second, on hardware built by the wafer-scale chipmaker Cerebras.
Ultrafast is not a new model so much as a new way to serve an existing one. It leans on Cerebras’s outsized chips to strip out the latency that has long dogged frontier AI, and OpenAI opened a limited preview on 13 August to a small group of customers, with plans to widen access as capacity allows.
The pitch turns on a trade-off OpenAI says it can finally dissolve. Until now, anyone who wanted genuinely real-time responses had to drop down to a smaller, less capable model, accepting less intelligence in exchange for speed.
Ultrafast is meant to deliver frontier-grade reasoning and near-instant answers at once, rather than forcing a choice between them.
That combination matters most for the agentic software the whole industry is chasing. An AI agent that has to think for thirty seconds before every step is a demo, whereas one that answers in the time it takes to hold a conversation starts to feel like a product.










