Papers for

automated math solving platforms

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multi-agent systems learn clearer roles and teamwork with on-policy distillation

MAS-OPD: On-Policy Distillation for Multi-agent Systems

Abstract: Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.

Mon 28 SeptComputation and Language
The gist
Splitting tasks among multiple specialized agents helps solve complex problems, but getting them to work well together is hard. The authors introduce MAS-OPD, a new way to teach these agents by giving detailed feedback from a teacher system during their learning. This approach helps each agent understand its role better and coordinates teamwork without needing complex, task-specific rewards. Tests on coding and math tasks show that MAS-OPD outperforms other methods and leads to better cooperation among agents.
Open → 2609.34234v1

Hidden state trajectories improve detection and repair of reasoning errors in diffusion language models

LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

Abstract: Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.

Wed 23 SeptComputation and Language
The gist
Sometimes, AI models that write answers get stuck on wrong ideas early and don’t change them later, even though they keep thinking. The authors studied how to spot and fix these mistakes by looking inside the model’s thinking process, not just the final answer or how confident it seems. They created a method called LOCKR that checks the model’s hidden thinking path and decides when to try other options to fix errors. This method improved the model’s accuracy on math reasoning tasks by a few percentage points and fixed many wrong answers.
Open → 2609.27220v1