OpenAI’s GPT-6 Astra scored 80% on a benchmark that asks vision-language models to find mistakes in partially assembled IKEA furniture. The result is a sharp improvement over the 28% achieved by the leading model in November 2025, although the test remains narrow and the system is not fast enough for real-time help.

Epoch AI’s Furniture Assembly Benchmark photographs three products at different stages while researchers deliberately introduce errors. A model must compare each image with the official instructions, locate the incorrect step and explain what went wrong. Claude Fable 5.1 scored 70%, while Claude Opus 5 reached 61%. The strongest Chinese open-weight systems trailed the leaders by at least seven months, according to the benchmark analysis.

Astra took about three minutes to process each photo. That delay limits immediate uses such as guiding someone while they work, and three furniture items cannot represent the full variety of physical troubleshooting. Even so, the jump shows measurable progress on detailed visual reasoning rather than general image description. If speed and reliability improve, the same approach could support appliance repair or vehicle maintenance where users must match a real object to a multi-step manual.