product updateOpenAI

GPT-6 Astra Ultrafast Claims Up to 8x Faster Token Generation on NVIDIA Blackwell GPUs

TL;DR

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. NVIDIA says it generates tokens up to 8x faster than Astra Standard mode. Pricing and context window details were not disclosed in the source.

2 min read
0

OpenAI's GPT-6 Astra Ultrafast is available now in the OpenAI API and to eligible ChatGPT Work and Codex users, running on NVIDIA Blackwell GPUs. According to NVIDIA, Ultrafast delivers up to 8x faster token generation than the Astra Standard mode.

What was announced

The details below come from NVIDIA's blog post, not from an independent test.

  • Availability: Live now in the OpenAI API and for eligible ChatGPT Work and Codex users. The source does not define "eligible."
  • Hardware: NVIDIA Blackwell GPUs.
  • Speed: Up to 8x faster token generation than Astra Standard mode. This is a vendor claim, and "up to" means the figure is a ceiling, not a typical result.
  • Method: NVIDIA attributes the gain to inference optimizations in OpenAI's models that use the capabilities of the Blackwell architecture.

Ultrafast is framed as a mode, not a new model

NVIDIA describes Ultrafast as a mode relative to "Astra Standard." It does not describe a separate trained model with its own weights. Based on the source, this reads as a serving configuration of GPT-6 Astra optimized for latency. We have not seen a distinct model identifier or a separate model card.

What is not disclosed

The source excerpt leaves out several figures developers will need:

  • Pricing per 1M tokens: not yet disclosed
  • Context window: not yet disclosed
  • Benchmark scores: not yet disclosed, including whether Ultrafast matches Standard on quality
  • Parameter count and training cutoff: not disclosed
  • Baseline throughput: the 8x figure has no stated tokens-per-second baseline, batch size, or prompt length

The source text was also truncated at the section on developer benefits, so further details may exist in the full post.

What this means

The headline number matters less than the missing conditions. An 8x gain over Standard is large, but without a baseline throughput, latency-to-first-token data, or quality comparison, builders cannot yet judge it. The key open questions are whether Ultrafast costs more per token than Standard and whether output quality is unchanged. Faster serving often comes with a price premium or a tradeoff, and neither has been confirmed here.

The release also shows how closely frontier model providers and hardware vendors are now aligned on inference. OpenAI names Blackwell as the platform for its fastest tier, and NVIDIA is publicizing the result. That points to inference speed becoming a competitive axis alongside model quality, particularly for agentic coding tools like Codex, where many sequential generation steps make latency the bottleneck.

Until OpenAI publishes its own documentation with pricing, limits, and measured performance, treat the 8x figure as a ceiling claimed by NVIDIA and test it against your own workloads.

Related Articles

product update

OpenAI launches Dots agent on GPT-6 Astra, limited to $20/month Pro tier and above

OpenAI unveiled Dots at DevDay 2026, a personal agent powered by GPT-6 Astra, available for now only on its $20/month Pro plan and above. Meta's rival Muse agent is free for anyone with a Meta account. OpenAI is positioning Dots as a long-horizon, knowledge-work tool with stronger privacy controls.

product update

OpenAI says it shut down a 15,000-account campaign to extract model reasoning; the attack still worked on Azure

OpenAI says it detected and shut down an adversarial distillation campaign involving more than 15,000 accounts, linked in part to people associated with Moonshot AI. Researchers report that the same reasoning-extraction attack still worked on Microsoft Azure on September 13 against every OpenAI model they tested, including GPT-6 Astra.

model release

OpenAI Ships GPT-6.1 Sol at DevDay 2026, Claims Near-Astra Performance at One-Fifth the Price

OpenAI's DevDay 2026 keynote introduced GPT-6.1 Sol, a mid-tier model priced at $2/$10 per million tokens that OpenAI claims delivers 'near-Astra intelligence' at a fraction of the cost. Independent benchmarks from Artificial Analysis and third-party testers show it trailing flagship Astra by roughly one point on the Intelligence Index while beating Opus 5.5 on cost-adjusted coding tasks.

model release

OpenAI's GPT-6.1 Sol Launches on Amazon Bedrock, Claims Near-Astra Reasoning at Fraction of Cost

OpenAI's GPT-6.1 Sol is now generally available on Amazon Bedrock, targeting agentic coding, computer use, and document-heavy business workflows. OpenAI claims the model matches GPT-6 Astra on the DeepSWE v1.1 coding benchmark at roughly one-fifth the cost per task.

Comments

Loading...