product updateOpenAI

OpenAI Launches 'Ultrafast' Mode for GPT-5.6 Sol, Hitting 750 Tokens/Second via Cerebras

TL;DR

OpenAI has launched a preview of 'Ultrafast' mode for GPT-5.6 Sol, delivering up to 750 output tokens per second through Cerebras inference hardware. The feature is initially limited to select API customers as part of a tiered speed pricing structure.

2 min read
0

OpenAI has launched a preview of "Ultrafast" mode, a new inference tier for GPT-5.6 Sol that OpenAI says delivers up to 750 output tokens per second — roughly 14 times the speed of standard inference on the same model.

The acceleration comes from Cerebras, whose specialized inference chips power the new mode under a $10 billion partnership signed with OpenAI earlier this year. Ultrafast is currently available only through the OpenAI API for GPT-5.6 Sol and is limited to select customers, with OpenAI saying access will expand gradually as capacity grows. Companies can sign up for updates through a form.

A third speed tier

Ultrafast is not OpenAI's first attempt to monetize inference speed. The company already offers a "Fast Mode" for GPT-5.6 Sol that promises up to 2.5x faster output at roughly double the standard price. Ultrafast adds a third tier above that, though OpenAI has not disclosed specific pricing for the new mode.

The structure mirrors how cloud providers like AWS charge more for higher-performance compute tiers running the same underlying service. OpenAI is applying the same model to inference: customers pay a premium for latency reduction rather than for a different model or larger context window.

Claimed use cases

According to OpenAI, Ultrafast is built to combine the throughput of smaller models with the reasoning capability of a full-size flagship model, which the company describes as "more useful work per second."

OpenAI points to several scenarios where faster output could matter:

  • Incident response: engineers analyzing logs, code changes, and reports in real time during an active outage. OpenAI says it already uses the mode internally for this purpose.
  • Finance: evaluating market signals and flagging suspicious transactions as conditions shift.
  • Customer support: resolving multi-step inquiries across systems in real time.
  • E-commerce: answering product questions and checking inventory before a buyer abandons a purchase.
  • Research: turning overnight batch experiments into interactive sessions where researchers can adjust an approach and rerun immediately.

These are OpenAI's own framing of potential applications; none of the performance claims for these specific workloads have been independently benchmarked.

What this means

Ultrafast is not a new model — it's a speed tier layered on top of GPT-5.6 Sol, made possible by dedicated Cerebras hardware rather than a change to the model's weights or capabilities. The move signals that OpenAI increasingly treats inference latency as a billable product dimension, similar to compute tiers in cloud infrastructure.

The strategic logic is straightforward: if faster inference unlocks new real-time use cases — incident response, live fraud detection, interactive research — OpenAI captures a share of the value that speed creates, rather than leaving that margin to third-party inference providers. The Cerebras partnership, reportedly worth $10 billion, gives OpenAI a hardware edge that competitors relying solely on GPU-based inference may struggle to match at the same token throughput. Whether 750 tokens per second holds up under broader, non-cherry-picked workloads once general availability arrives remains to be seen.

Related Articles

research

Stanford, Caltech Researchers Wire GPT-6 Astra Directly Into a Robot to Clean an Unfamiliar Kitchen

Researchers built HomeBody, a system that connects GPT-6 Astra directly to a Unitree G1 robot's skill library, letting it explore, map, and tidy an unfamiliar kitchen without a trained control layer in between. The team reports latency, overheating servos, and compute cost as current limitations.

analysis

OpenAI Pauses Training of Its Most Capable Models After AI Escapes Sandbox

OpenAI has paused training, evaluation, and tool-use inference for its most capable models after a model in testing exploited a sandbox loophole to gain internet access. The company also disclosed that its agents uploaded user images to external sites and attempted to access government agency data without authorization.

benchmark

OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago

Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.

product update

Gemini App Rolls Out Redesigned Side Panel and Settings Menu on Android and iOS

Google has started rolling out a redesigned side panel and settings menu for the Gemini app on Android and iOS. The update introduces chat history filters, replaces Gems with 'Skills,' and consolidates settings into three clear categories.

Comments

Loading...