A new arXiv position paper argues that AI researchers need more precise definitions of reasoning before they can reliably measure progress. The authors say current generative AI work often treats reasoning as an intuitive capability rather than an operational construct.

Their proposal is to reconnect reasoning evaluation with rule-based and verifiable traditions from symbolic AI. That does not mean abandoning neural models. It means defining tasks, rules and success conditions clearly enough that benchmark results can be checked for construct validity.

The paper matters because reasoning has become a marketing and research label for many model releases. Without shared definitions, two systems can be compared on tests that appear similar while measuring different abilities.

This is a position paper, so it offers a framework rather than a finished benchmark. Its useful limit is also its point: claims about autonomous reasoning should be tied to explicit rules and evidence, not just fluent explanations or high scores on loosely defined tasks.