product updateOpenAI

OpenAI Launches 'Ultrafast' Mode for GPT-5.6 Sol, Hitting 750 Tokens/Second via Cerebras

TL;DR

OpenAI has launched a preview of 'Ultrafast' mode for GPT-5.6 Sol, delivering up to 750 output tokens per second through Cerebras inference hardware. The feature is initially limited to select API customers as part of a tiered speed pricing structure.

2 min read
0

OpenAI has launched a preview of "Ultrafast" mode, a new inference tier for GPT-5.6 Sol that OpenAI says delivers up to 750 output tokens per second — roughly 14 times the speed of standard inference on the same model.

The acceleration comes from Cerebras, whose specialized inference chips power the new mode under a $10 billion partnership signed with OpenAI earlier this year. Ultrafast is currently available only through the OpenAI API for GPT-5.6 Sol and is limited to select customers, with OpenAI saying access will expand gradually as capacity grows. Companies can sign up for updates through a form.

A third speed tier

Ultrafast is not OpenAI's first attempt to monetize inference speed. The company already offers a "Fast Mode" for GPT-5.6 Sol that promises up to 2.5x faster output at roughly double the standard price. Ultrafast adds a third tier above that, though OpenAI has not disclosed specific pricing for the new mode.

The structure mirrors how cloud providers like AWS charge more for higher-performance compute tiers running the same underlying service. OpenAI is applying the same model to inference: customers pay a premium for latency reduction rather than for a different model or larger context window.

Claimed use cases

According to OpenAI, Ultrafast is built to combine the throughput of smaller models with the reasoning capability of a full-size flagship model, which the company describes as "more useful work per second."

OpenAI points to several scenarios where faster output could matter:

  • Incident response: engineers analyzing logs, code changes, and reports in real time during an active outage. OpenAI says it already uses the mode internally for this purpose.
  • Finance: evaluating market signals and flagging suspicious transactions as conditions shift.
  • Customer support: resolving multi-step inquiries across systems in real time.
  • E-commerce: answering product questions and checking inventory before a buyer abandons a purchase.
  • Research: turning overnight batch experiments into interactive sessions where researchers can adjust an approach and rerun immediately.

These are OpenAI's own framing of potential applications; none of the performance claims for these specific workloads have been independently benchmarked.

What this means

Ultrafast is not a new model — it's a speed tier layered on top of GPT-5.6 Sol, made possible by dedicated Cerebras hardware rather than a change to the model's weights or capabilities. The move signals that OpenAI increasingly treats inference latency as a billable product dimension, similar to compute tiers in cloud infrastructure.

The strategic logic is straightforward: if faster inference unlocks new real-time use cases — incident response, live fraud detection, interactive research — OpenAI captures a share of the value that speed creates, rather than leaving that margin to third-party inference providers. The Cerebras partnership, reportedly worth $10 billion, gives OpenAI a hardware edge that competitors relying solely on GPU-based inference may struggle to match at the same token throughput. Whether 750 tokens per second holds up under broader, non-cherry-picked workloads once general availability arrives remains to be seen.

Related Articles

changelog

OpenAI Launches 'Ultrafast' Mode, Claims 14x Speed Boost for GPT 5.6 Sol via Cerebras Partnership

OpenAI has introduced 'Ultrafast,' a preview mode that it claims accelerates GPT 5.6 Sol to 14 times standard speed, hitting up to 750 output tokens per second. The feature runs on OpenAI's partnership with chipmaker Cerebras and is currently limited to a small group of customers.

changelog

OpenAI Previews 'Ultrafast' Tier for GPT-5.6 Sol, Claims Up to 14x Speed Increase

OpenAI is testing an 'Ultrafast' service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second using Cerebras infrastructure. Access is currently limited to a waitlist of select customers.

product update

OpenAI Launches Computer History: A Local, Searchable Timeline of macOS Activity for ChatGPT Memory

OpenAI has launched Computer History, a macOS feature that records clicks, keystrokes, and app switches to build a searchable memory timeline for ChatGPT and Codex. It replaces the screenshot-based Chronicle preview and requires opt-in consent from both admins and individual users.

product update

OpenAI Replaces Chronicle with Computer History in ChatGPT for Mac

OpenAI has replaced ChatGPT for Mac's Chronicle research preview with Computer History, an opt-in feature that builds a searchable activity timeline using macOS accessibility events instead of screenshots. The feature is available to ChatGPT Pro, Business, and Enterprise users, excluding the EEA, Switzerland, and the UK.

Comments

Loading...