OpenAI is launching a preview of its new "Ultrafast" mode, which delivers up to 750 output tokens per second from its flagship model, GPT-5.6 Sol.

The inference acceleration comes from Cerebras, which signed a ten-billion-dollar partnership with OpenAI earlier this year. The service will initially be available only through the OpenAI API for GPT-5.6 Sol and limited to select customers. OpenAI plans to expand access gradually as capacity grows. Companies that want in can sign up for updates through a form.

Faster output opens up new use cases

OpenAI says Ultrafast is designed to combine the speed of smaller models with the full capabilities of a large reasoning model, enabling what the company calls "more useful work per second." During incident response, for example, engineers could have logs, code changes, and reports analyzed while an outage is still happening, helping them pinpoint the cause and prepare a fix in real time. OpenAI says it's already using the model internally for this purpose.

OpenAI pitches several other scenarios. In finance, the model could evaluate market signals and flag suspicious transactions while conditions are still shifting. In customer support, complex inquiries could be resolved in real time, even when finding the answer requires multiple steps or systems. In e-commerce, it could answer product questions, check inventory levels, and personalize recommendations before a hesitant buyer abandons their cart.