benchmarkOpenAI

Unverified 'GPT-6 Astra' Reportedly Completes Portal Solo in Under 24 Hours, No Official OpenAI Confirmation

TL;DR

A developer named cozyblaze posted on X that a model called 'GPT-6 Astra' completed Portal from start to finish without human intervention in 23 hours 43 minutes. OpenAI has not confirmed the existence of a model by that name, and all details come from a single third-party account.

2 min read
0

What was reported

A developer using the handle cozyblaze posted on X claiming that a model called "GPT-6 Astra" played through the entirety of Portal and reached the game's credits without human help after an initial goal was set. According to the post, the run took approximately 23 hours and 43 minutes.

No model by the name "GPT-6 Astra" has been announced or confirmed by OpenAI. This claim originates entirely from a third-party social media post and an accompanying GitHub repository, not from any official OpenAI communication, blog post, or API documentation.

Claimed technical details

According to cozyblaze, the model controlled Portal through the Model Context Protocol (MCP) combined with a modified tool called SourcePauseTool, which pauses the game engine while the model processes information. During each pause, the agent reportedly received screenshots, player position data, and camera angle information, selected inputs, then allowed the game to resume. A published video of the run reportedly has these pause intervals edited out.

Cozyblaze claims token usage for the run cost at least $570 at what is described as "Astra's list price," though the developer says a $200 Codex subscription was used instead. Neither the existence of an "Astra" pricing tier nor these cost figures can be independently verified — no such model or pricing appears in OpenAI's public documentation as of this writing.

Context from the source

Cozyblaze referenced a 2016 OpenAI goal of building a single agent capable of solving many different games, framing this Portal run as a partial realization of that vision. The developer also stated that this version is "the worst model we'll ever get" — a subjective claim with no independent support.

What this means

Treat this report as unconfirmed. The claim rests entirely on one developer's social media post and self-published code repository; there is no statement from OpenAI acknowledging a model called GPT-6 Astra, no benchmark data from a neutral third party, and no verifiable pricing information. Names like "GPT-6" carry significant weight, and unofficial leaks or fabricated model names circulate frequently in AI communities, sometimes to generate attention rather than convey verified capability.

If accurate, an agent autonomously completing a full 3D puzzle-platformer using only screenshots and spatial reasoning over a 24-hour span would represent a meaningful demonstration of long-horizon planning and visual grounding. But without confirmation from OpenAI — including basic facts like the model's actual existence, its context window, or its real pricing — this should be read as an interesting but unverified community claim, not a confirmed model release or benchmark result. Readers should wait for official acknowledgment before treating any of these figures as fact.

Related Articles

model release

OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing

A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.

model release

OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol

OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.

benchmark

Simon Willison's Pelican Benchmark Shows GPT-6 Astra Outperforming GPT-5.6 Sol at Every Reasoning Level

Developer Simon Willison ran his signature 'pelican riding a bicycle' SVG test on newly-accessed GPT-6 Astra across five reasoning levels, comparing results against GPT-5.6 Sol, Terra, and Luna. Even Astra's lowest reasoning setting reportedly beat every Sol output, though Astra costs roughly twice as much per token.

product update

OpenAI Lists GPT-6 Astra Pro on OpenRouter: Same Model, Higher-Compute Reasoning Mode

GPT-6 Astra Pro, now listed on OpenRouter, is the existing GPT-6 Astra model configured to run with reasoning.mode set to 'pro' for higher-quality output on complex tasks. It carries a 1M-token context window and tiered pricing from $5/$25 to $20/$100 per million input/output tokens depending on the serving tier.

Comments

Loading...