OpenAI is previewing Ultrafast, a new API tier that runs GPT-5.6 Sol at up to 14 times the speed of standard processing, delivering up to 750 output tokens per second1,2. The tier is powered by chipmaker Cerebras.
The preview was announced on August 13. OpenAI described Ultrafast as a step toward delivering more useful work per second from its most capable model, rather than requiring users to trade down to smaller or more specialized models for real-time speed. "Until now, getting real-time speed typically meant choosing a smaller or more specialized model," the company said in a blog post. "Ultrafast points to progress in a new direction: more useful work per second".
The preview is initially available to a small group of customers, with OpenAI saying it will expand access as capacity grows. OpenAI suggested Ultrafast can be deployed across corporate workflows including incident response, customer service and support, financial market analysis, and e-commerce.
The Cerebras partnership underpins Ultrafast as an infrastructure choice. ANALYSIS The collaboration allows OpenAI to offer throughput levels on GPT-5.6 Sol that exceed what its standard serving stack delivers, based on the 14x speed differential the company cited.
TechCrunch noted that competitors have pursued similar acceleration strategies: Anthropic's Claude offers a "fast mode," though TechCrunch reported it does not deliver the speed levels OpenAI is claiming for Ultrafast.
ANALYSIS The limited preview rollout to a small group of customers, with plans to expand as capacity grows, suggests capacity constraints on the Cerebras-powered infrastructure.
The announcement lands on the same day OpenAI named Dali Rajic, president and COO of Wiz, as its new chief revenue officer ctx. ◆ An enterprise-facing inference product launching alongside a CRO transition aligns with OpenAI's commercial positioning of its model capabilities.