OpenAI just made its most powerful model a lot faster. The company announced a limited preview of “Ultrafast mode” for GPT-5.6 Sol on August 13, delivering processing speeds up to 14 times faster than the standard version of the same model. The target audience: enterprise customers who need AI that can keep up with real-time workflows like voice applications, financial research, and security response.
The key number is 750 output tokens per second. For context, that’s roughly the equivalent of generating an entire page of text in about one second, fast enough that the bottleneck in most applications shifts from “waiting for the AI” to “figuring out what to do with the answer.”
What Cerebras brings to the table
The speed gains aren’t coming from software tricks alone. OpenAI partnered with Cerebras, the AI chip company known for building wafer-scale processors that dwarf conventional GPUs, to power the Ultrafast tier. Cerebras hardware enables GPT-5.6 Sol to run 11 times faster than Fable 5, OpenAI’s previous-generation model, and 5 times faster than Opus 4.8 running on Fast mode.
Why 14x speed matters for enterprise AI










