OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago
Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.
OpenAI's GPT-6 Astra scored 80% on Epoch AI's Furniture Assembly Benchmark (FAB), a test that measures whether AI models can spot mistakes in IKEA furniture builds by comparing progress photos against assembly instructions. Ten months earlier, in November 2025, the best-performing model on the same benchmark—Claude Opus 4.5—scored just 28%.
FAB works by photographing three IKEA furniture pieces at various stages of assembly, with deliberate errors introduced during the build. Models are shown the photos alongside the official instructions and asked to identify what went wrong and describe the mistake. It's a test of fine-grained visual reasoning: matching physical object states against a sequence of diagrams, then localizing discrepancies.
According to Epoch AI, Claude Fable 5.1 placed second at 70%, followed by Claude Opus 5 at 61%. Chinese open-weight models, including Kimi K3, trail the leading closed models by at least seven months on this benchmark, according to the researchers.
GPT-6 Astra's 80% score comes with a catch: it takes roughly three minutes per photo to process, according to Epoch AI. That processing time rules out real-time use during an actual furniture build—by the time the model finishes analyzing a single image, a person could likely have already found their own mistake. But the researchers suggest the underlying capability could eventually extend to other structured visual-inspection tasks, such as diagnosing car repairs or troubleshooting appliance faults.
Epoch AI frames the score jump as notable given how recently multimodal models were failing much simpler visual tasks. A near-tripling of accuracy in ten months on a benchmark specifically designed to test detailed visual-instruction matching represents one of the sharper capability jumps tracked on any narrow visual benchmark this year. The report also notes that GPT-6 Astra performs strongly on visual robotic tasks more broadly, suggesting the gains extend beyond the furniture-specific test.
No pricing, context window, or architectural details for GPT-6 Astra were disclosed in the source material. Epoch AI's FAB benchmark itself remains a narrow, purpose-built test rather than a general capability suite, so the 80% figure should be read as a specific measure of visual-error-detection performance rather than a broad claim about model quality.
What this means
Fine-grained visual comparison—matching real-world object states against reference instructions and pinpointing exact discrepancies—has historically been a weak spot for multimodal models, which tend to describe images in broad strokes rather than catch small structural errors. A jump from 28% to 80% in under a year suggests this specific capability is improving faster than general benchmark trends would predict. If the three-minute-per-image latency comes down, the practical applications extend well beyond furniture: quality control in manufacturing, remote technical support, and any workflow where someone needs to verify that a physical process matches a documented procedure. The gap between closed frontier models and open-weight alternatives on this task—cited at seven-plus months—also signals that detailed visual reasoning remains one of the areas where proprietary labs are extending their lead rather than seeing it close.
Related Articles
Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands
A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.
OpenAI Launches Astra for Law, a Legal Research Tool Built on GPT-6 Astra
OpenAI has launched Astra for Law, a legal-focused version of GPT-6 Astra that combines the model with a case law search index and specialized analysis instructions. The tool scored 54 percent on Vals AI's Legal Research Bench in OpenAI's own testing, up from 38.7 percent for the base model with web search.
OpenAI Gives ChatGPT Voice Access to Email, Calendar, and Slack, Powered by New GPT-6 Models
OpenAI has rolled out a major ChatGPT Voice upgrade that lets users manage email, calendar events, and Slack messages by voice. The feature now runs on new GPT-6 Astra, Sol, and Luna models and is available globally in the latest app version.
OpenAI Upgrades ChatGPT Voice with GPT-6 Power, Plugin Support, and ChatGPT Work Integration
OpenAI is upgrading ChatGPT Voice with three changes: it now runs on GPT-6 models, supports plugins like email and Slack, and integrates with ChatGPT Work on web and mobile. The update addresses a longstanding gap between voice mode and OpenAI's broader feature set.
Comments
Loading...