Papers for

software engineers building virtual tutors

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Planned test-time scaling improves reasoning task performance with coordinated problem solving

Planned Test-Time Scaling with Coordinated Reasoning Paths

Abstract: Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each branch, steering the branches toward distinct reasoning paths, and an executor produces a full solution conditioned on each outline. Formally, we show that PTTS strictly generalizes repeated sampling and, in a stylized setting, provably promotes coverage of complementary reasoning modes and yields better pass@k scaling. We instantiate PTTS on top of strong reasoning models, keeping them fixed as executors while replacing repeated sampling with PTTS inference to further enhance test-time scaling. Concretely, we develop two variants: PTTS-ZS prompts a model to jointly generate outlines for all branches in a single autoregressive pass, while PTTS-RL directly optimizes the planner against the pass@k reward using truncated execution rollouts for efficient training and a sharper reward signal. Across five mathematical reasoning benchmarks with Qwen3-1.7B and 4B, PTTS-ZS improves pass@64 over repeated sampling by up to 6.7 points, while PTTS-RL further increases the gain to up to 13.4 points. Further analysis indicates that broader coverage of distinct reasoning paths contributes to these gains. Overall, PTTS provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.

Wed 23 SeptComputation and LanguageArtificial Intelligence
The gist
When computers try to solve hard problems, they often try many guesses at once to find answers. The authors show that making these guesses work together, instead of independently, helps the computer explore different kinds of solutions better. They created a new method called Planned Test-Time Scaling (PTTS) that first plans different outlines for each guess and then completes each one. This approach improves success rates on math reasoning tests by encouraging diverse ways of thinking. It works both without extra training and with training to get even better results.
Open → 2609.27374v1