OpenAI's Ultrafast Mode Pushes GPT-5.6 Sol to 750 Tokens a Second

The Core · TL;DR
- OpenAI's new Ultrafast mode runs GPT-5.6 Sol at 14x normal speed, hitting up to 750 tokens per second
- The speedup comes from a partnership with Cerebras, using its Wafer-Scale Engine chip with 44 GB of integrated SRAM
- On the 'Humanity's Last Exam' benchmark, Sol Ultrafast finished 2,500 questions in 11h11m versus 78h27m for Fable 5
- Artificial Analysis data shows Sol Ultrafast is about 11x faster than Fable 5 and 5x faster than Anthropic's Opus 4.8 Fast mode; access is currently limited to a small preview group
OpenAI is now running GPT-5.6 Sol at 14 times its normal processing speed through a new mode called Ultrafast, announced on August 13, 2026. The feature pushes output to as much as 750 tokens per second, a pace that turns tasks once measured in hours into ones measured in minutes.
The gain comes from hardware, not just software. Ultrafast runs on Cerebras' Wafer-Scale Engine, a chip architecture built around 44 GB of integrated SRAM memory that keeps data close to compute rather than shuttling it across a network of GPUs.
That architecture explains the scale of the speedup better than any prompt-engineering trick could. Cerebras has spent years positioning wafer-scale silicon as an inference accelerator, and OpenAI's adoption of it for a flagship model is a concrete bet that the approach can now serve production traffic.
The benchmark numbers make the case vivid. Running the 2,500-question "Humanity's Last Exam" test, Sol Ultrafast finished in 11 hours and 11 minutes, compared with 78 hours and 27 minutes for OpenAI's own Fable 5.
Independent tracking from Artificial Analysis puts Sol Ultrafast at roughly 11 times faster than Fable 5 and about 5 times faster than Anthropic's Opus 4.8 running in its own Fast mode. That comparison places OpenAI's latency advantage squarely against its closest commercial rival rather than against its own prior generation alone.
Where the speed matters
OpenAI is pitching Ultrafast for jobs where waiting costs money or safety margin: incident response, financial market analysis, customer service, e-commerce, voice applications, coding, and research workflows. In each case, the value isn't smarter answers, it's faster ones, delivered while the underlying situation is still current.
For a trading desk or an on-call engineer, a model that replies in fractions of a second changes what's practically usable in the loop, not just what's theoretically possible.
Access remains limited for now. Ultrafast is in preview and rolled out to a small group of customers, with OpenAI saying it will widen availability as Cerebras capacity grows.
That caveat matters for anyone evaluating the mode today. The headline speed figures describe what the system can do under current conditions, not what every developer can expect to reserve immediately, and broader rollout will depend on how quickly wafer-scale inference capacity can be added.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
