model releaseOpenAI

OpenAI Launches GPT-6 Astra, Claims SOTA Computer Use and Coding — But Independent Tests Show Mixed Gains at Higher Cost

TL;DR

OpenAI released GPT-6 Astra on September 3, 2026, claiming state-of-the-art computer use and coding performance alongside new alignment techniques. Independent evaluators found real but uneven gains, higher per-task costs, and reduced chain-of-thought monitorability.

3 min read
0

OpenAI released GPT-6 Astra on September 3, 2026, calling it the company's "most intelligent and aligned model yet." The launch drew 36 million views and 164,000 likes within nine hours, according to OpenAI's own social metrics — the company's largest reaction since Sora, surpassing the receptions for GPT-4 and GPT-5.

Pricing and Availability

Astra ships in two tiers. Standard pricing is $10 per 1M input tokens and $50 per 1M output tokens. A "fast" tier costs $20 per 1M input tokens and $100 per 1M output tokens in exchange for roughly 2.5x higher throughput. That represents a 2.5x price increase per token over OpenAI's prior flagship, GPT-5.6 Sol.

Rollout began with a limited set of organizations, then expands to ChatGPT Plus, Pro, Business, and Enterprise tiers, followed by API and AWS access over the following days, according to OpenAI. The staggered release caused friction: paying ChatGPT users reported delayed access while select influencers received early access, prompting OpenAI to issue "banked resets" — extra usage credit for each day of lost access — to affected subscribers.

Claimed Benchmarks

OpenAI's own figures include 99.9% on ARC-AGI-3, 98% on FrontierMath Tier 4, 100% on ExploitBench, and a claim that Astra runs 1.9x faster than GPT-5.6 Sol on Mind2Web with Codex harness improvements. OpenAI also stated Astra "already helped solve long-standing open problems in mathematics," citing work related to prime-gap research — a claim not yet independently verified.

Independent Results Tell a More Mixed Story

Artificial Analysis measured Astra at 67 on its Coding Agent Index — roughly matching Claude Opus 5 and Fable 5, but trailing Fable 5.1's score of 70. On the Intelligence Index, Astra scored 61, tying GPT-5.6 Sol and trailing Claude Fable 5.1 by 5 points. Token efficiency improved substantially: Astra used one-third the tokens of GPT-5.6 Sol and one-fifth the tokens of Claude Opus 5 for comparable coding tasks. However, at maximum reasoning effort, the 2.5x price increase made Astra roughly 75% more expensive per task than its predecessor, despite using about 10% fewer output tokens.

Hallucination rates reportedly dropped from 92% to 51% at max effort on Artificial Analysis's benchmark, with a 4-point accuracy gain. Long-horizon knowledge work improved by roughly 80 Elo on the AA-Briefcase benchmark, though Presentation Quality Elo declined versus GPT-5.6 Sol. Notable regressions appeared on GDPval-AA v2 (down ~80 Elo) and 2-3 point drops on τ³-Banking, SciCode, and AA-LCR.

ARC Prize evaluators reported divergent scores depending on harness: 63% under Astra's direct scoring versus 99% using a new provider adapter harness. François Chollet independently measured 66% on the standard harness, near 100% using a continuous conversation harness with custom compaction, at a cost of roughly $360 per game.

Alignment and Monitorability Concerns

OpenAI's accompanying system card described improved alignment metrics alongside decreased chain-of-thought monitorability — meaning Astra's internal reasoning traces are harder for external observers to audit for safety issues. Researchers including Neel Nanda and Ryan Greenblatt raised concerns that visible alignment improvements may mask rather than resolve underlying goal-misalignment failure modes, rather than eliminate them.

What This Means

GPT-6 Astra delivers genuine efficiency gains — using a fraction of the tokens of rival models for comparable coding output — but the 2.5x price hike largely offsets those savings at high reasoning effort, undercutting OpenAI's headline claims of dramatic capability jumps. The reduced chain-of-thought monitorability is the more consequential development for the field: as models get better at solving benchmarks, verifying why they reach an answer is becoming harder, not easier. Expect competitive responses from Anthropic, Google DeepMind, and xAI within weeks, with monitorability likely becoming a differentiator alongside raw benchmark scores.

Related Articles

model release

OpenAI Launches GPT-6 Astra, Claims State-of-the-Art Computer Use and 98% on FrontierMath Tier 4

OpenAI has launched GPT-6 Astra, claiming state-of-the-art results on computer use, coding, and scientific reasoning benchmarks. The model is rolling out to a limited set of organizations first, with general ChatGPT availability expected within days.

model release

OpenAI Launches GPT-6 Astra, Matches Claude Fable Pricing at $10/$50 per Million Tokens

OpenAI has begun rolling out GPT-6 Astra, priced at $10/million input and $50/million output tokens to match Claude Fable. The model claims a 99.9% score on ARC-AGI 3 using a custom harness and leads on security and long-context benchmarks, though it trails Fable on general intelligence rankings.

model release

OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold

OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.

model release

OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing

A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.

Comments

Loading...