changelogOpenAI

OpenAI Launches 'Ultrafast' Mode, Claims 14x Speed Boost for GPT 5.6 Sol via Cerebras Partnership

TL;DR

OpenAI has introduced 'Ultrafast,' a preview mode that it claims accelerates GPT 5.6 Sol to 14 times standard speed, hitting up to 750 output tokens per second. The feature runs on OpenAI's partnership with chipmaker Cerebras and is currently limited to a small group of customers.

2 min read
0

OpenAI has rolled out a new processing mode called Ultrafast for GPT 5.6 Sol, its top-tier model, claiming it delivers output at 14 times the speed of standard processing — up to 750 output tokens per second, according to the company.

The feature was announced in a blog post published Thursday and is currently available only in preview to a limited set of customers. OpenAI says it will widen access "as capacity grows," though no timeline has been given.

How it works

Ultrafast is powered by OpenAI's partnership with chipmaker Cerebras, whose wafer-scale processors are built for high-throughput inference. OpenAI frames the mode as a shift in approach: rather than trading model capability for speed by routing to a smaller model, Ultrafast aims to keep the power of GPT 5.6 Sol while cutting latency.

"Until now, getting real-time speed typically meant choosing a smaller or more specialized model," OpenAI said in its announcement. "Ultrafast points to progress in a new direction: more useful work per second."

No pricing has been disclosed for Ultrafast access, and OpenAI has not published independent benchmark data verifying the 14x speed claim or the 750 tokens/second figure — both numbers currently rest solely on the company's own statement.

Target use cases

OpenAI is pitching Ultrafast at latency-sensitive enterprise workflows, specifically:

  • Incident response
  • Customer service and support
  • Financial market analysis
  • E-commerce

These are all scenarios where response time directly affects user experience or operational outcomes, making raw inference speed a more visible differentiator than raw capability.

Competitive context

Anthropic already offers a "fast mode" for Claude, but according to the TechCrunch report, it does not match the throughput OpenAI is claiming for Ultrafast. Speed has become a growing axis of competition among frontier labs alongside benchmark performance and context window size, as more enterprise deployments move from experimentation into production systems where latency has direct cost and revenue implications.

What this means

Ultrafast is not a new model — it's a serving mode for the existing GPT 5.6 Sol checkpoint, similar to how Anthropic offers a fast variant of Claude. The real story here is infrastructure: OpenAI leaning on Cerebras's specialized hardware to squeeze more throughput out of a large model without shrinking it down.

If the 750 tokens/second figure holds up under independent testing, it would meaningfully close the gap between frontier-model quality and the sub-second responsiveness that real-time applications like voice agents and live customer support demand. But until OpenAI publishes verifiable benchmarks or opens Ultrafast beyond a small preview group, the 14x claim remains an unverified performance target rather than a documented result. Enterprises evaluating latency-critical deployments should treat this as promising but unproven until broader access — and independent testing — becomes available.

Related Articles

benchmark

Simon Willison's Pelican Benchmark Shows GPT-6 Astra Outperforming GPT-5.6 Sol at Every Reasoning Level

Developer Simon Willison ran his signature 'pelican riding a bicycle' SVG test on newly-accessed GPT-6 Astra across five reasoning levels, comparing results against GPT-5.6 Sol, Terra, and Luna. Even Astra's lowest reasoning setting reportedly beat every Sol output, though Astra costs roughly twice as much per token.

model release

OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold

OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.

product update

Pentagon Adds OpenAI's ChatGPT Mil and xAI's Grok for Government to GenAI.mil

The Pentagon has added OpenAI's ChatGPT Mil and xAI's Grok for Government to its GenAI.mil platform, which previously offered only Google Gemini. Anthropic's Claude remains excluded after a supply-chain risk dispute with the Trump administration.

product update

OpenAI Lists GPT-6 Astra Pro on OpenRouter: Same Model, Higher-Compute Reasoning Mode

GPT-6 Astra Pro, now listed on OpenRouter, is the existing GPT-6 Astra model configured to run with reasoning.mode set to 'pro' for higher-quality output on complex tasks. It carries a 1M-token context window and tiered pricing from $5/$25 to $20/$100 per million input/output tokens depending on the serving tier.

Comments

Loading...