OpenAI Launches Ultrafast Mode — GPT-5.6 Sol Runs 14x Faster on Cerebras

OpenAI partners with Cerebras to deliver GPT-5.6 Sol at 750 tokens per second — 14x faster than standard inference — enabling real-time AI applications.

Thursday August 13, 2026 Source: OpenAI
TL;DR — Quick Answer

OpenAI launched 'ultrafast mode' running GPT-5.6 Sol on Cerebras wafer-scale hardware at roughly 750 output tokens per second — about 14x faster than standard deployment. The deep integration makes real-time voice assistants, live coding copilots, and interactive agents viable at near-human conversational speed.

Key Takeaways

OpenAI Launches Ultrafast Mode — GPT-5.6 Sol Runs 14x Faster on Cerebras — AI news article illustration

OpenAI has announced a new “ultrafast mode” that runs GPT-5.6 Sol on Cerebras hardware, delivering inference speeds roughly 14 times faster than standard deployment — with output tokens streaming at approximately 750 tokens per second.

The Hardware-Software Stack

The ultrafast mode is the result of a deep integration between OpenAI’s model architecture and Cerebras’ wafer-scale chip technology. Unlike conventional GPU-based inference, Cerebras’ approach processes entire layers of a neural network on a single massive chip, eliminating the inter-chip communication overhead that typically bottlenecks speed.

Performance Numbers

The speed gains are achieved through the combination of Cerebras’ chip architecture and OpenAI’s inference optimizations, which have been co-developed specifically for this deployment.

What It Means for Users

For developers and businesses relying on OpenAI’s API, the ultrafast mode has immediate practical implications:

The Cerebras Partnership

Cerebras has positioned itself as the alternative to NVIDIA’s dominance in AI chips, and this OpenAI partnership represents a significant validation. By running one of OpenAI’s most capable models on non-NVIDIA hardware, Cerebras demonstrates that the AI chip market is not a monopoly.

For OpenAI, the partnership provides a hedge against NVIDIA supply constraints and pricing power, while also delivering genuine performance improvements that benefit customers.

Availability

The ultrafast mode is available in preview via the OpenAI API, with broader rollout expected in the coming weeks. Developers can opt in by specifying the Cerebras-backed inference tier in their API calls.

Frequently Asked Questions

What is OpenAI ultrafast mode?

Ultrafast mode is an OpenAI API option that runs GPT-5.6 Sol on Cerebras wafer-scale hardware, delivering inference at roughly 750 output tokens per second — about 14x faster than standard deployment — with sub-100ms time-to-first-token.

How fast is GPT-5.6 Sol on Cerebras?

The ultrafast mode streams approximately 750 output tokens per second, which is roughly 14x faster than OpenAI's standard GPT-5.6 Sol endpoint, with sub-100ms time-to-first-token for most prompts.

Does ultrafast mode reduce model quality?

No. Ultrafast mode delivers the full GPT-5.6 Sol capability with no reduction in model quality or context window — the speed gains come from Cerebras' wafer-scale chip architecture eliminating inter-chip communication overhead.

This article is based on the official announcement from OpenAI . Read the original for full technical details.

Related Articles

Back to all news