OpenAI has announced a new “ultrafast mode” that runs GPT-5.6 Sol on Cerebras hardware, delivering inference speeds roughly 14 times faster than standard deployment — with output tokens streaming at approximately 750 tokens per second.
The Hardware-Software Stack
The ultrafast mode is the result of a deep integration between OpenAI’s model architecture and Cerebras’ wafer-scale chip technology. Unlike conventional GPU-based inference, Cerebras’ approach processes entire layers of a neural network on a single massive chip, eliminating the inter-chip communication overhead that typically bottlenecks speed.
Performance Numbers
- 750 output tokens per second — roughly 14x faster than OpenAI’s standard GPT-5.6 Sol endpoint
- Sub-100ms time-to-first-token for most prompts
- Full GPT-5.6 Sol capability — no reduction in model quality or context window
The speed gains are achieved through the combination of Cerebras’ chip architecture and OpenAI’s inference optimizations, which have been co-developed specifically for this deployment.
What It Means for Users
For developers and businesses relying on OpenAI’s API, the ultrafast mode has immediate practical implications:
- Real-time applications become far more viable — voice assistants, live coding copilots, and interactive agents can respond at near-human conversational speed
- Batch processing becomes dramatically faster and cheaper
- User experience improves — lower latency means higher perceived quality, even at the same model capability
The Cerebras Partnership
Cerebras has positioned itself as the alternative to NVIDIA’s dominance in AI chips, and this OpenAI partnership represents a significant validation. By running one of OpenAI’s most capable models on non-NVIDIA hardware, Cerebras demonstrates that the AI chip market is not a monopoly.
For OpenAI, the partnership provides a hedge against NVIDIA supply constraints and pricing power, while also delivering genuine performance improvements that benefit customers.
Availability
The ultrafast mode is available in preview via the OpenAI API, with broader rollout expected in the coming weeks. Developers can opt in by specifying the Cerebras-backed inference tier in their API calls.