OpenAI previewed Ultrafast on August 13, a new tier for GPT-5.6 Sol that runs up to 14 times faster than standard processing, powered by Cerebras. It generates up to 750 output tokens per second, fast enough to sit inside real-time products instead of behind an overnight batch job. Sol is the flagship of the GPT-5.6 family that launched in July, alongside the balanced Terra and cost-efficient Luna. The speed comes from Cerebras’s wafer-scale chips, which pack 44 GB of memory onto a single wafer so model weights stay on-chip and never travel to memory. Ultrafast is a limited preview in the OpenAI API for a select group of customers including Jane Street, Podium, Basis, and Rogo, with access expanding as capacity grows.
Until now, real-time speed meant choosing a smaller or more specialized model. Ultrafast removes that tradeoff for time-sensitive work: incident response while an outage is unfolding, financial research while markets move, customer support, and research loops that used to run overnight. Cerebras says its output is 11 times faster than Claude Fable 5 and 5 times faster than Opus 4.8 on Fast mode. In its own test of Humanity’s Last Exam, a 2,500-question PhD-level benchmark, Sol Ultrafast finished in 11 hours and 11 minutes where Fable 5 needed 78 hours and 27 minutes, about 7 times longer. Pricing is not yet public, and these figures are vendor-run.
Cerebras chips sit inside OpenAI’s API, speed is sold as a tier, and the flagship, not a smaller sibling, is the fast one.
Read More: GPT-5.6 Sol’s Cycle Double Cover Proof: What Lean Actually Checks
Sources:
- Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed (OpenAI)
- Accelerating GPT-5.6 Sol Ultrafast with OpenAI (Cerebras)
- Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier (Unite.AI)
- GPT-5.6: Frontier intelligence that scales with your ambition (OpenAI)
Disclaimer: For information only. Accuracy or completeness not guaranteed. Illegal use prohibited. Not professional advice or solicitation. Read more: /terms-of-service
Reuse
Citation
@misc{kabui2026,
author = {{Kabui, Charles}},
title = {GPT-5.6 {Sol} {Ultrafast:} {OpenAI’s} {New} {Tier} {Runs} 750
{Tokens} a {Second}},
date = {2026-08-24},
url = {https://toknow.ai/posts/gpt-5-6-sol-ultrafast-cerebras-750-tokens-per-second/},
langid = {en-GB}
}
