AI summaryⓘ
The authors study how to best use limited thinking resources when solving multiple problems together, instead of giving each problem its own separate resources. They create a new test called R³-Bench that looks at performance on six problems under shared budgets in math, programming, and reasoning tasks. Their findings show that an ideal offline strategy usually does better than actual tested models, highlighting a gap between what models can do individually and how they manage resources across tasks. They also find that simple equal allocation strategies sometimes work better than contest approaches, but no single method is best across all areas. This points to challenges in optimizing resource use across multiple problems simultaneously.
resource rationalityshared budgetscognitive sciencemulti-task allocationR³-Benchoffline oracleagentic pressuresingle-problem competencestrategy updatingtrajectory diagnostics
Authors
Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
Abstract
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.