Anthropic’s Claude Opus 5 has opened a large lead on ARC-AGI-3, a benchmark built to test how AI systems handle unfamiliar interactive tasks rather than memorized examples. ARC Prize reports that the model scored 30.2 percent, compared with the previous top score of 7.8 percent from OpenAI’s GPT-5.6 Sol Max.

The benchmark asks a model to infer the rules of a new environment, plan actions, and execute them step by step. ARC Prize says Opus 5 solved five environments that had not been solved before and showed new behaviors during testing, including translating tasks into algebraic notation and formulating reflection equations.

That makes the result notable for researchers tracking model reasoning, but it is not a general intelligence finish line. Six of the 25 public demo environments have now been solved, which means most remain beyond current systems.

The older ARC-AGI-1 and ARC-AGI-2 scores are also high for Opus 5, though ARC Prize says some results came at higher cost. The useful signal is narrower: stronger planning and exploration are showing up in tests designed to be less like standard question answering.