GPT-6 Astra fully completed seven of 100 physical tasks in StationeryBench, a new robotics test involving ordinary desk objects. Ai2’s MolmoAct2 completed none of its 100 attempts while controlling the same dual-arm YAM robots.
The benchmark covers five tasks such as uncapping a marker, pouring paper clips and passing a ruler between two arms. Beyond the small number of complete successes, Astra earned a median progress score of 46 out of 100, compared with 12 for MolmoAct2. The researchers have published code, results and videos, making the comparison open to inspection.
A separate, still-unpublished benchmark called REMAP reportedly places Astra close to human accuracy on some spatial questions. Researcher Yoav Artzi described the improvement as a step change, while emphasizing that the model still falls short of people in other scenarios. He speculated that extensive three-dimensional training data, such as Blender scenes, could explain the gains, but that has not been confirmed. Seven successful trials out of 100 remain a clear practical limitation: the result is evidence of progress in spatial understanding, not a demonstration that general-purpose robots can yet depend on the model.