GPT-6 Astra Beats Ai2's MolmoAct2 on New Robotics Benchmark, Researcher Calls It a 'Step Change'
A new robotics benchmark called StationeryBench shows OpenAI's GPT-6 Astra completing 7 of 100 desk-object manipulation tasks versus zero for Ai2's MolmoAct2, with a median progress score of 46 against 12. Cornell/DeepMind researcher Yoav Artzi calls the result a 'step change in spatial reasoning.'
OpenAI's GPT-6 Astra has posted a substantial lead over Ai2's MolmoAct2 on a new robotics benchmark designed to test spatial reasoning in physical manipulation tasks, according to results published on GitHub and reported by Robocurve.
The benchmark
The test, called StationeryBench, pits the two models against each other across five desk-object tasks: uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms, among others. Both models controlled identical dual-arm YAM robots across 200 trials total.
Astra fully completed 7 out of 100 tasks; MolmoAct2 completed zero. On a graded progress metric, Astra's median score reached 46 out of 100, compared to 12 for MolmoAct2. All results, videos, and code are publicly available on GitHub.
What researchers are saying
Yoav Artzi, an AI researcher affiliated with Cornell and Google DeepMind, described Astra's performance as a "step change in spatial reasoning." On a separate, still-unpublished benchmark called REMAP, Artzi says GPT-Astra reaches accuracy close to human level — though he cautions that "even ASTRA doesn't get to what humans do in other scenarios."
Artzi speculates that OpenAI trained the model on large volumes of 3D data, possibly including Blender-rendered scenes, which he says would explain Astra's particular strength on 3D spatial tasks relative to prior models.
Neither OpenAI nor Ai2 has published a formal technical report accompanying these results, and the benchmark itself — StationeryBench — is new and has not undergone independent peer review or widespread replication. The REMAP benchmark cited by Artzi remains unpublished entirely, meaning its methodology and scoring have not been made available for outside scrutiny.
OpenAI has previously disclosed long-term plans to build its own consumer robots, and Astra's spatial reasoning capabilities are widely seen as a building block toward that goal, though the company has not confirmed a product timeline.
What this means
A 46 versus 12 median progress score and a 7-versus-0 full-completion gap are meaningful margins, but the sample size — 200 trials across five tasks on one robot platform — is narrow, and neither StationeryBench nor REMAP has been independently validated. Robotic manipulation benchmarks are also notoriously sensitive to hardware setup, task design, and grading rubrics, so these results should be read as an early, single-source signal rather than a settled measure of spatial reasoning ability.
Still, the direction is consistent with a broader industry push to close the gap between language-model reasoning and physical-world competence — the bottleneck that has kept general-purpose robotics out of reach. If OpenAI's 3D-data training hypothesis holds up under scrutiny, it suggests a relatively straightforward path (more 3D-rendered training data) toward better spatial models, which would be significant for any company building robotics or embodied AI products, not just OpenAI.
Related Articles
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price
An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.
OpenAI Launches ChatGPT for Financial Services to Automate Wall Street Analyst Work
OpenAI launched ChatGPT for Financial Services, a tailored enterprise product built with design partners Morgan Stanley and Evercore that automates research, financial analysis, and pitchbook creation. The tool, powered by GPT-6 Astra, targets tasks traditionally performed by Wall Street's junior analysts and associates.
Perplexity Says It Runs End-to-End Engineering Systems on OpenAI's GPT-6 Astra
Perplexity says it has shifted core engineering workflows, including code changes and production monitoring, onto OpenAI's GPT-6 Astra model. The claim comes from an OpenAI-published case study with no independent benchmark data released.
Comments
Loading...