OpenAI's GPT-6 Astra Dominates IKEA Assembly Test
OpenAI's GPT-6 Astra has achieved an 80 percent accuracy rate on Epoch AI's Furniture Assembly Benchmark, signaling a massive leap forward in how AI models interpret complex physical tasks.

OpenAI's GPT-6 Astra has set a new standard for computer vision by scoring 80 percent on Epoch AI's Furniture Assembly Benchmark (FAB). The benchmark evaluates an AI's ability to identify physical mistakes by analyzing photographs of three different IKEA furniture pieces taken during assembly. These photos contain deliberate errors, requiring models to compare the images against official instructions, pinpoint the exact mistakes, and describe what went wrong.
The rapid progress highlights a massive leap in spatial reasoning. In November 2025, the top-performing model on the FAB was Anthropic's Claude Opus 4.5, which scored a mere 28 percent. Just ten months later, GPT-6 Astra reached its 80 percent high-water mark, though it currently requires three minutes of processing time per photo. Other frontier models also showed significant improvement, with Claude Fable 5.1 reaching 70 percent and Claude Opus 5 scoring 61 percent. Meanwhile, Chinese open-weight models like Kimi K3 continue to lag behind the industry leaders by at least seven months.
For AI practitioners and robotics engineers, this rapid advancement represents a major shift in how multimodal models interact with the physical world. While a processing time of three minutes per photo makes the technology too slow for real-time, interactive assembly assistance today, the underlying capabilities are highly promising. Researchers suggest that this level of visual understanding will eventually translate to practical applications like automated car repairs, home appliance troubleshooting, and advanced visual robotic tasks, where Astra already shows strong performance.
This is our own summary of reporting by The Decoder



