OpenAI on August 13 previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 14 times the speed of standard processing, generating up to 750 output tokens per second1,2,4. The tier is powered by chipmaker Cerebras.
Ultrafast is launching first in the OpenAI API and is currently available only in preview to a small group of customers. OpenAI said it will expand access as capacity grows.
"Until now, getting real-time speed typically meant choosing a smaller or more specialized model," OpenAI said in its blog post. "Ultrafast points to progress in a new direction: more useful work per second".
The company positioned the tier for latency-sensitive enterprise workflows including incident response, customer service, financial market analysis, and e-commerce.
A separate Reddit post claims OpenAI acquired a 4.2% stake in Cerebras ahead of the Ultrafast launch3. That claim is unverified and carries no corroborating sourcing beyond the social-media post.
TechCrunch noted that Anthropic has launched a "fast mode" for Claude, though it characterized OpenAI's speed claims as exceeding what that competitor offers.
Vector Wire covered the initial Ultrafast announcement on August 14[1]. The new detail surfacing since that coverage is the unverified Cerebras equity claim.
ANALYSIS The Cerebras partnership is the infrastructure story beneath the speed headline. By offloading inference to Cerebras hardware, OpenAI is decoupling its fastest inference tier from its own GPU fleet — a departure from the vertically integrated serving model that has defined most frontier-lab deployments. If the reported 4.2% equity stake is accurate, it would indicate a financial relationship that extends beyond a standard cloud-compute contract, though the claim remains unverified.
The 750-token-per-second output rate at full model capability — rather than on a distilled or smaller variant — targets a segment of enterprise demand where latency has historically forced customers to trade off model quality for speed. The explicit naming of incident response and financial-market use cases signals that OpenAI is marketing Ultrafast as production infrastructure, not a research preview.