GPT-6 Astra tops Ai2's MolmoAct2 in new spatial-reasoning benchmark

GPT-6 Astra tops Ai2's MolmoAct2 in new spatial-reasoning benchmark

A new robotics benchmark called StationeryBench pitted OpenAI's GPT-6 Astra against Ai2's MolmoAct2 on five desk-object task types, among them uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms. Both models controlled the same dual-arm YAM robots across 200 trials.

Astra fully completed 7 of the 100 tasks it was scored on, while MolmoAct2 completed none. Across the trials, Astra's median progress score reached 46 out of 100, against 12 for MolmoAct2. StationeryBench's results, videos and code are posted on GitHub, linked directly from the report. Separately, OpenAI is said to have long-term plans to build its own consumer robots, with no timeline or product details disclosed.

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, called the result "a step change in spatial reasoning." He also pointed to a separate, still-unpublished benchmark called REMAP, on which the report says GPT-Astra reaches accuracy close to human level, while Artzi cautioned that "even ASTRA doesn't get to what humans do in other scenarios." He suspects OpenAI trained the model on large amounts of 3D data, such as scenes built in Blender, a theory that the report says would line up with Astra's particular improvement on 3D tasks.

Key facts

  • StationeryBench pitted OpenAI's GPT-6 Astra against Ai2's MolmoAct2 on five desk-object task types, including uncapping a marker and pouring out paper clips, using the same dual-arm YAM robots across 200 trials.
  • Astra fully completed 7 of 100 tasks; MolmoAct2 completed zero.
  • Astra's median progress score reached 46 out of 100, against 12 for MolmoAct2.
  • Yoav Artzi, an AI researcher at Cornell and Google DeepMind, called the gap "a step change in spatial reasoning" and suspects Astra was trained on large amounts of 3D data such as Blender scenes.
  • On the still-unpublished REMAP benchmark, the report says GPT-Astra reaches accuracy close to human level, though Artzi cautions it still falls short of humans in other scenarios.

Why it matters

Most claims about a frontier lab's spatial reasoning gains come from the lab itself. Here an academic outside both companies, Yoav Artzi of Cornell and Google DeepMind, called the gap between GPT-6 Astra and Ai2's MolmoAct2 "a step change." Robot manipulation has been a stubborn weak spot for language-model-derived systems, so a documented jump on a controlled, five-task benchmark carries more weight than another leaderboard score on a text or coding test.

Who it affects

OpenAI's own robotics ambitions most directly: the report says OpenAI has long-term plans to build its own consumer robots, and manipulation tasks like these are the building blocks such robots would need. It also affects Ai2, whose MolmoAct2 lost on both metrics tested, and researchers building or evaluating vision-language-action models for warehouse, home or desk-based manipulation, who now have a new benchmark and a specific gap to try to close.

How to use it

StationeryBench's results, videos and code are posted on GitHub, linked directly from the report, so the underlying material is at least nominally open to outside inspection rather than described only secondhand. The report says nothing about pricing, access or a release timeline for GPT-6 Astra itself.

How solid is it

The comparison ran on physical hardware, the same dual-arm YAM robots for both models, across 200 trials, rather than in simulation, and the endorsement comes from an academic outside both companies rather than from OpenAI's own marketing. Two things temper that: the report does not say who built StationeryBench or the REMAP benchmark, so it is unclear whether either test is independent of OpenAI, and it does not explain how the five task types and 200 trials map onto the 100-task tallies behind the completion and progress-score numbers. The REMAP claim itself carries no published number, only the qualitative phrase "close to human level."

Risks and caveats

Even the winning score is modest in absolute terms: Astra fully completed just 7 of 100 tasks, and a median progress score of 46 out of 100 means most trials were only partly finished. The report does not quantify Astra's "particular improvement on 3D tasks" separately from its overall scores, so the size of that specific effect is unclear, and Artzi's theory that OpenAI trained on Blender-style 3D data is his own supposition, not something OpenAI has confirmed. The article itself refers to the model as both "GPT-6 Astra" and "GPT-Astra" without clarifying whether these name the same system, and OpenAI's own consumer-robot plans come with no timeline or product details.

“step change in spatial reasoning”

— Yoav Artzi, AI researcher at Cornell and Google DeepMind