Improving large language model reasoning with efficient step compression

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

Computation and LanguageArtificial IntelligenceMachine Learning

Summary

Large Language Models (LLMs) get better at reasoning when they think through problems step-by-step, but writing out all those steps can take a lot of time and computer power. The authors developed a new method called A*-Thought-V2 that smartly compresses less important steps into shorter hidden code while keeping key steps in clear text. This saves resources and makes the model faster without losing accuracy. They also introduced ways to train the model to balance detailed thinking and efficient compression. Experiments show this approach improves accuracy and speeds up processing across various tasks.

Large Language ModelsChain-of-Thoughtlatent representationdimensionality reductionprincipal component analysismodel compressionreasoning dynamicssoft labelingtrajectory modeling

Authors

Xiaoang Xu, Siyuan Liu, Shuo Wang, Junlan Feng, Fanyu Meng, Zhu Zhang, Jixun Wang, Xiaorong Wang, Zihan Zhou, Xin Li, Chaojun Xiao, Yiming Zhang, Huijia Wu, Liuyu Xiang, Peipei Li, Zhaofeng He

Abstract

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29$\times$, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.