OpenAI Launches 'Ultrafast' Mode, Claims 14x Speed Boost for GPT 5.6 Sol via Cerebras Partnership
OpenAI has introduced 'Ultrafast,' a preview mode that it claims accelerates GPT 5.6 Sol to 14 times standard speed, hitting up to 750 output tokens per second. The feature runs on OpenAI's partnership with chipmaker Cerebras and is currently limited to a small group of customers.
OpenAI has rolled out a new processing mode called Ultrafast for GPT 5.6 Sol, its top-tier model, claiming it delivers output at 14 times the speed of standard processing — up to 750 output tokens per second, according to the company.
The feature was announced in a blog post published Thursday and is currently available only in preview to a limited set of customers. OpenAI says it will widen access "as capacity grows," though no timeline has been given.
How it works
Ultrafast is powered by OpenAI's partnership with chipmaker Cerebras, whose wafer-scale processors are built for high-throughput inference. OpenAI frames the mode as a shift in approach: rather than trading model capability for speed by routing to a smaller model, Ultrafast aims to keep the power of GPT 5.6 Sol while cutting latency.
"Until now, getting real-time speed typically meant choosing a smaller or more specialized model," OpenAI said in its announcement. "Ultrafast points to progress in a new direction: more useful work per second."
No pricing has been disclosed for Ultrafast access, and OpenAI has not published independent benchmark data verifying the 14x speed claim or the 750 tokens/second figure — both numbers currently rest solely on the company's own statement.
Target use cases
OpenAI is pitching Ultrafast at latency-sensitive enterprise workflows, specifically:
- Incident response
- Customer service and support
- Financial market analysis
- E-commerce
These are all scenarios where response time directly affects user experience or operational outcomes, making raw inference speed a more visible differentiator than raw capability.
Competitive context
Anthropic already offers a "fast mode" for Claude, but according to the TechCrunch report, it does not match the throughput OpenAI is claiming for Ultrafast. Speed has become a growing axis of competition among frontier labs alongside benchmark performance and context window size, as more enterprise deployments move from experimentation into production systems where latency has direct cost and revenue implications.
What this means
Ultrafast is not a new model — it's a serving mode for the existing GPT 5.6 Sol checkpoint, similar to how Anthropic offers a fast variant of Claude. The real story here is infrastructure: OpenAI leaning on Cerebras's specialized hardware to squeeze more throughput out of a large model without shrinking it down.
If the 750 tokens/second figure holds up under independent testing, it would meaningfully close the gap between frontier-model quality and the sub-second responsiveness that real-time applications like voice agents and live customer support demand. But until OpenAI publishes verifiable benchmarks or opens Ultrafast beyond a small preview group, the 14x claim remains an unverified performance target rather than a documented result. Enterprises evaluating latency-critical deployments should treat this as promising but unproven until broader access — and independent testing — becomes available.
Related Articles
OpenAI Gives ChatGPT Voice Access to Email, Calendar, and Slack, Powered by New GPT-6 Models
OpenAI has rolled out a major ChatGPT Voice upgrade that lets users manage email, calendar events, and Slack messages by voice. The feature now runs on new GPT-6 Astra, Sol, and Luna models and is available globally in the latest app version.
OpenAI Upgrades ChatGPT Voice with GPT-6 Power, Plugin Support, and ChatGPT Work Integration
OpenAI is upgrading ChatGPT Voice with three changes: it now runs on GPT-6 models, supports plugins like email and Slack, and integrates with ChatGPT Work on web and mobile. The update addresses a longstanding gap between voice mode and OpenAI's broader feature set.
OpenAI Launches GPT-6 Sol and GPT-6 Luna, Cutting API Prices 50% Versus GPT-5.6
OpenAI has released GPT-6 Sol and GPT-6 Luna, two new models that cost 50% less than their GPT-5.6 equivalents while claiming improved coding and computer-use performance. The models roll out today to ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.
OpenAI's GPT-6 Sol and GPT-6 Luna Launch on Amazon Bedrock, Priced Below GPT-5.6
OpenAI has launched GPT-6 Sol and GPT-6 Luna on Amazon Bedrock, positioned below flagship GPT-6 Astra for recurring coding tasks and high-volume document processing respectively. Both models cost less per API call than their GPT-5.6 predecessors, though exact pricing figures were not disclosed.
Comments
Loading...