OpenAI is launching "Ultrafast," a new inference mode that delivers GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from their $10 billion partnership. Together with "Standard" and "Fast," Ultrafast creates a three-tier pricing structure that turns inference speed into its own product.

OpenAI is previewing a new way to run its most capable GPT-5.6 model at dramatically higher speeds. The company says...

Google's Gemini 3.7 Flash model is live and built for cheap agents; OpenAI's GPT-5.6 Sol Ultrafast is quicker but locked behind a waitlist.