A new arXiv paper proposes structural uncertainty as a way to measure consistency in LLM logical reasoning. Instead of only checking whether sampled final answers differ, the method asks whether a model can consistently rank competing reasoning solutions.

The authors argue that models may reach the same answer through unstable or contradictory reasoning paths. That failure mode is easy to miss if evaluation focuses only on output dispersion.

The framework is relevant for high-stakes reasoning systems where a correct answer is less reassuring if the model’s path to it is inconsistent.