Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

2026-08-17Artificial Intelligence

Artificial Intelligence
AI summary

The authors argue that mathematical creativity is made up of several different ways of creating meaning in math, such as thinking deeply about math practice, using ideas from science, solving specific problems, and connecting far-apart ideas. They say that current large language models mainly work by mixing and reusing existing patterns, which means they might not be able to do all these types of creativity. Because proving math statements is getting easier with AI, the value in math is shifting to types of creativity that current models don’t handle well. Therefore, the authors suggest we should evaluate AI math ability using this breakdown instead of broad tests that mix everything together.

mathematical inventionmeaning-makingconjecture formationtransformer modelsmathematical creativityanalogical reasoningAI evaluationbenchmarkingproblem-driven constructioncross-domain bridging
Authors
Silvère Gangloff
Abstract
How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - reflexive introspection on mathematical practice, analogical import from the sciences, problem-driven construction, and the bridging of distant domains - together with a further, cross-cutting distinction between meaning pursued because a pattern was observed and meaning pursued because it is strategically wanted, a distinction I develop through the case of conjecture-formation. These mechanisms are likely non-substitutable, so that competence in one does not transfer to the others. Grounding each in a historical case study and in an architecture-level account of current transformer-based systems, I suggest that today's models concentrate their competence in modes shaped by recombination and search over existing building blocks; if that description holds, the remaining modes are out of reach in principle, not just slower - though whether it holds is itself the open, empirical part. Because proof is getting cheaper as AI improves at generating it - a shift the field's own leading voices are now diagnosing - mathematical value is migrating toward the modes current systems cannot yet perform, and evaluations of AI mathematical ability should be organized around this taxonomy rather than around aggregate benchmarks that conflate it.