Large language models improved by fixing math reasoning steps

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Machine Learning

Summary

Solving math problems with AI is tricky because it needs clear step-by-step understanding, not just guesses. The authors studied what kinds of math reasoning large language models (LLMs) can and cannot do well, finding that discovering new math ideas is the hardest part. They created a test to measure different reasoning skills and found ways to fix common mistakes after training. Their new method helps models learn better math reasoning by focusing on these key reasoning parts, improving their problem-solving skills.

What this means in practice

  • For ai developers: Enhance AI systems' ability to solve complex math problems by incorporating targeted training of math reasoning primitives.
  • For software test engineers: Use the proposed benchmark to evaluate and diagnose mathematical reasoning strengths and weaknesses in AI models before deployment.

Authors

Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu

Abstract

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.