OpenAI Launches GPT-6 Astra, Says the Model May Already Qualify as AGI
OpenAI has released GPT-6 Astra, its most capable model yet, with benchmark scores the company says surpass GPT-5.6 Sol and Anthropic's Fable 5 models. President Greg Brockman called it a step into the 'AGI era,' though OpenAI acknowledges there's no agreed-upon threshold for that term.
GPT-6 Astra — Quick Specs
OpenAI Ships GPT-6 Astra, Its Most Capable Model to Date
OpenAI has launched GPT-6 Astra, a model the company says outperforms its predecessor GPT-5.6 Sol and Anthropic's Fable 5 models across reasoning, math, coding, and cybersecurity benchmarks. OpenAI President Greg Brockman said the model might already qualify as AGI — an AI system that outperforms humans at most economically valuable work, by OpenAI's own definition — or is close to it.
Astra is rolling out first to select organizations through OpenAI's Daybreak program. ChatGPT Plus, Pro, Business, and Enterprise customers will get access over the coming days, along with API users and cloud platforms including AWS Bedrock and Microsoft Azure.
Training and Benchmark Claims
According to OpenAI researcher Aidan Clark, Astra was pretrained on more than 100,000 GPUs at the company's Stargate facility in Texas, making it OpenAI's largest training run to date. Clark said the capability jump from Sol to Astra is larger than the jump that produced Sol, attributing part of the gain to earlier models assisting in monitoring the training process.
OpenAI's published benchmarks show Astra scoring 98.6 percent on ARC-AGI-3 (under OpenAI's own test conditions), 97.6 percent on FrontierMath Tier 4 v2, 96 percent on GPQA Diamond, 74.1 percent on DeepSWE v1.1, and a perfect 100 percent on ExploitBench, a cybersecurity benchmark measuring vulnerability discovery and exploitation. On long-context tests, Astra scored 96.3 percent on MRCR v2's 8-needle task at the 512K–1M token range, versus 73.8 percent for Sol at the same range — suggesting a context window of at least 1 million tokens, though OpenAI has not stated an exact figure.
On computer-use tasks, OpenAI says Astra scored 72.6 percent on OSWorld 2.0 while completing tasks in roughly 40 minutes on average, compared to Sol's 65.7 percent and 75 minutes. The company also claims Astra improved a decades-old mathematical result on prime gaps and set new records on internal biology, chemistry, medicine, and physics evaluations — claims that have not been independently verified.
Pricing: Higher Per Token, Lower Per Task, OpenAI Claims
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens in standard API mode — 2.5 times more than GPT-5.6 Sol and roughly in line with Anthropic's Fable 5.1. A faster mode, promising 2.5x speed, doubles that price again.
Brockman argued that per-token pricing is becoming a poor comparison metric since tokenization differs across model families. He said OpenAI is shifting toward pricing based on cost per completed task, and cited an internal estimate that Astra's top configuration cuts API costs per task by about 57 percent versus Sol on the DeepSWE v1.1 benchmark. This figure comes from OpenAI and has not been independently confirmed.
First Model Classified as "Critical" Risk
OpenAI says Astra is the first model to trigger a "critical" classification under its Preparedness Framework, meaning it can autonomously discover unknown vulnerabilities and build exploit chains against well-defended systems without step-by-step human guidance. The company says Astra found two previously unknown zero-day vulnerabilities during evaluation, which were reported to affected vendors and confirmed by human experts. OpenAI also acknowledges that Astra's internal reasoning is harder to monitor than Sol's. The most advanced cybersecurity capabilities remain restricted to trusted defenders through the Daybreak Blue program, and OpenAI says it delayed Astra's release to conduct additional safety testing.
During the briefing, Brockman conceded there is no universally agreed-upon threshold for AGI, noting the transition has been more gradual than OpenAI originally anticipated when the company was founded. He nonetheless closed by declaring, "Welcome to the AGI era." CEO Sam Altman had previously said he expected a model he would call AGI by the end of 2026.
What This Means
Astra's benchmark gains, if verified independently, represent a meaningful jump over Sol and Anthropic's Fable line, particularly in cybersecurity and computer-use tasks. But the "AGI" framing is a marketing and definitional call by OpenAI, not a scientific consensus — no external body has validated the claim, and Brockman himself admits there's no agreed threshold. The critical-risk cybersecurity classification is the more concrete and consequential development here: a model that can autonomously find and chain zero-day exploits raises real dual-use concerns independent of whether it meets anyone's definition of AGI.
Related Articles
OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks
OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.
OpenAI Launches GPT-6 Astra, Matches Claude Fable Pricing at $10/$50 per Million Tokens
OpenAI has begun rolling out GPT-6 Astra, priced at $10/million input and $50/million output tokens to match Claude Fable. The model claims a 99.9% score on ARC-AGI 3 using a custom harness and leads on security and long-context benchmarks, though it trails Fable on general intelligence rankings.
Meta's Muse Spark 1.3 Claims #3 Global Ranking, Matches OpenAI's GPT-5.6-Sol on Coding Benchmarks
Meta Superintelligence Labs shipped Muse Spark 1.3, which the company claims ranks #3 globally on the Artificial Analysis Intelligence Index and matches OpenAI's GPT-5.6-Sol on coding and agentic benchmarks. The model is available now via Muse Code and Meta's API, with open weights and a follow-up model promised soon.
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
Comments
Loading...