OpenAI's GPT-6 Astra hits 80% on IKEA assembly error benchmark

OpenAI's GPT-6 Astra hits 80% on IKEA assembly error benchmark

Epoch AI runs a test called the Furniture Assembly Benchmark (FAB), which photographs three IKEA pieces mid assembly with deliberate errors built in. Models are shown the photos alongside the instructions and have to compare the two, spot the mistake, and describe what went wrong. In November 2025 the best model on this test, Claude Opus 4.5, scored 28%. Ten months later, OpenAI's GPT-6 Astra reached 80%, taking about three minutes to work through each photo. Claude Fable 5.1 came in at 70% and Claude Opus 5 at 61%, while Chinese open-weight models such as Kimi K3 trail the leaders by at least seven months on this benchmark. The jump is notable because, as the report on the results points out, models were struggling with far simpler visual tasks not long ago. Astra is also described as excelling at visual robotic tasks more broadly. The system is still too slow to help with assembly in real time, but researchers say the underlying capability could eventually be applied to tasks like car repairs or appliance fixes, where spotting what went wrong from a photo against a reference is a similar problem.

Key facts

  • GPT-6 Astra scored 80% on Epoch AI's Furniture Assembly Benchmark, up from a best score of 28% by Claude Opus 4.5 in November 2025, ten months earlier
  • Claude Fable 5.1 scored 70% and Claude Opus 5 scored 61% on the same benchmark
  • Chinese open-weight models such as Kimi K3 trail the leading models by at least seven months on this test
  • GPT-6 Astra takes about three minutes per photo, still too slow to help with assembly in real time
  • Researchers say the approach could eventually assist with tasks like car repairs or appliance fixes

Why it matters

The benchmark tracks a narrow but concrete visual reasoning skill: comparing a photo of a real object against instructions and catching a deliberate mistake. Scores moving from 28% to 80% in ten months is a sharp jump on a task where models were failing much simpler visual checks not long ago, and it is being read as a marker of how fast visual reasoning is improving.

Who it affects

The comparison mainly matters to people tracking frontier model capability, since it stacks GPT-6 Astra against Claude Opus 4.5, Claude Fable 5.1, Claude Opus 5, and open-weight Chinese models like Kimi K3. It also matters to anyone imagining future use cases for this kind of photo-based error checking, such as home repair or appliance troubleshooting.

How to use it

There is no consumer product here yet. GPT-6 Astra needs about three minutes per photo, which the source notes is still too slow for real-time help while someone is mid-assembly. The researchers behind the work say the same photo-versus-instructions comparison could eventually be pointed at car repairs or appliance fixes, but that is described as a future direction, not something available now.

How solid is it

The figures come from Epoch AI's own Furniture Assembly Benchmark, with the 28% and 80% scores, the three-minute processing time, and the comparison scores for Claude Fable 5.1 and Claude Opus 5 stated directly. The exact date GPT-6 Astra reached 80% is not given, only that it came ten months after the November 2025 result, and no researcher is quoted by name.

Risks and caveats

The benchmark covers just three IKEA pieces with built-in errors, a narrow domain that may not generalize to other kinds of visual error-spotting. The source does not explain the mechanism by which the models identify mistakes, does not give a specific score for Kimi K3 beyond it trailing the leaders, and does not detail the full test methodology.