ARC-AGI-3
5 articles tagged with ARC-AGI-3
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
OpenAI Launches GPT-6 Astra, Claims State-of-the-Art Computer Use and 98% on FrontierMath Tier 4
OpenAI has launched GPT-6 Astra, claiming state-of-the-art results on computer use, coding, and scientific reasoning benchmarks. The model is rolling out to a limited set of organizations first, with general ChatGPT availability expected within days.
Nvidia Research: Custom Harness Pushes Claude Opus 5 from 30% to 100% on ARC-AGI-3
Nvidia researchers report that a custom agent harness with a supervisory component drove Claude Opus 5 from a 30% baseline score to a perfect 100% on the ARC-AGI-3 benchmark. The findings add to mounting evidence that scaffolding, not the underlying model, is the dominant factor in long-horizon agentic performance.
OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
Claude Opus 5 Scores 30.2% on ARC-AGI-3, Nearly 4x the Previous Record
Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the leap to genuinely stronger reasoning, though an independent test on a separate puzzle benchmark showed far smaller improvements.