KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
2026-07-10 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors identified that current process reward models (PRMs), which help improve large language model (LLM) multi-agent systems, are slow because they re-read all the text from scratch, especially for long tasks. They created a new model called KV-PRM that uses an existing internal memory structure (KV cache) from the LLM to score results much faster. This new approach reduces the computing time and memory needed by a huge margin while still performing as well or better than traditional text-based methods on several benchmarks. The authors also proved that the KV cache holds more useful information than just the text alone, making their method more efficient for reward modeling.
Process Reward ModelsTest-Time ScalingLarge Language ModelsKV CacheMulti-Agent SystemsBeam SearchMonte Carlo Tree SearchFLOPsLatencyMemory Footprint
Authors
Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang
Abstract
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000x reduction in scoring FLOPs, a 37x reduction in latency, and a 34x reduction in per-sequence memory footprint compared to text-based PRMs.