Selective inference improves tree structured decision making in reinforcement learning

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Artificial Intelligence

Summary

When computers try to make decisions by exploring different possible future moves, they often pick choices based on past results. But this can cause a problem because some options look better just because they were selected before, not necessarily because they truly are better. The authors found a way to correct for this bias by carefully re-evaluating choices and their alternatives, making the decision-making process fairer and more accurate. Their method, called SIPO, showed better performance in question-answering tests using large language models.

What this means in practice

  • For machine learning engineers: Improve decision-making algorithms that use tree search by correcting bias from repeated selections in rollout data.
  • For qa system developers: Enhance multi-step question-answering models by better credit assignment across alternative answer continuations.

Authors

Zenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han, Zhaoyang Li, Qiuyuan Ai, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang

Abstract

Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by $1.31$ and $1.07$ percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO