Two prominent mathematicians say large language models have become useful mathematical calculators, but remain weak at the kind of creative abstraction behind major discoveries.
Timothy Gowers argues that current models are good at combining known methods and exploring many search paths. The limitation is choosing the few promising paths in an enormous problem space, a skill mathematicians often describe as intuition.
Peter Sarnak makes a similar distinction. In his view, AI can derive results from existing theory but does not yet develop the new abstractions that make difficult proofs possible when the starting point is a simple question.
The assessment aligns with DeepMind researcher Tom Zahavy’s paper “LLMs Can’t Jump,” which identifies a bottleneck in “manipulative abduction,” or inventing new foundational assumptions without a clear linguistic precedent. The practical message is not that AI is useless for math. It is that benchmark progress and theorem manipulation do not automatically mean models can originate the conceptual leaps that define new fields.