Moonshot AI’s PerceptionBench suggests that frontier multimodal models still have serious visual perception limits, according to The Decoder. The benchmark tries to isolate whether models can correctly read an image before judging their reasoning about it.
That separation matters because many failures in vision-language systems are described as reasoning mistakes. If the model misreads an object, chart or spatial relation at the first step, a more elaborate chain of thought cannot fix the answer.
The report says no frontier model reached 60% accuracy, with GPT-5.6 Sol leading by a narrow margin. That leaves a large gap between impressive demo behavior and dependable visual understanding.
Benchmarks are only one lens, and performance can vary by task type. Still, PerceptionBench points to a practical limit for agents that operate on screenshots, documents, interfaces or real-world images: seeing the input correctly remains a bottleneck.