Language models teach themselves to solve math problems better

Teach to Learn: Hint Annealing for Self-improving LLM Reasoning

Machine Learning

Summary

Sometimes language models struggle to learn from difficult math problems because their initial answers are wrong, so they get no helpful feedback. The authors found that giving the model hints helps it learn, but relying on hints too much can hurt its ability to solve problems independently. They designed a method called HATCH that lets the model create and use its own hints while gradually reducing dependence on them, improving its reasoning skills over time without external help. Tests showed this approach worked better than previous ones on math reasoning tasks.

What this means in practice

  • For software engineers: Improve AI systems that automatically solve or check math problems by enabling models to self-teach and reason more accurately.
  • For ai chatbot developers: Enhance chatbot reasoning capabilities to provide better step-by-step explanations and answers without relying on external hints.

Authors

Zile Wang, Zijian Li, Haodong Wang, Jian Liu, Qianli Liu, Lucas Muli, Blaze Chen, Song Guo

Abstract

Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based methods construct auxiliary hints from solution evidence and use them to re-solve failed queries, recovering learning signal. Yet the resulting trajectories are typically treated as ordinary solution trajectories despite being generated under an assisted condition unavailable at evaluation. We discover hinted reward shift: recovered reward contrast can concentrate policy updates on hinted trajectories, limiting improvement without hints. This also creates a trade-off: increasing hinted trajectories can accelerate early learning but intensify reward shift later. To address this problem, we propose HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance. To mitigate hinted reward shift, we introduce online weighting to anneal the contribution of hinted trajectories. However, learning to generate hints can conflict with improving query solving. We therefore use gradient projection to remove the opposing component of hint-generation updates. Together, these designs support self-improvement by enabling the policy to create learning opportunities for itself and turn them into stronger reasoning without hints. We evaluate our method on mathematical reasoning benchmarks and outperform state-of-the-art methods by 1.02 pp on Llama-3.2-1B-Instruct, 2.84 pp on Qwen3-1.7B, and 4.32 pp on Qwen3-8B.