CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
2026-08-06 • Machine Learning
Machine LearningComputation and Language
AI summaryⓘ
The authors present CalibForge, a system that automatically creates tasks for training AI agents by testing and adjusting these tasks based on how different solvers perform on them. They use two methods to measure solver performance: one that looks for disagreements among various solvers, and another that checks if a strong solver can solve a task while a weak one cannot. By doing this, CalibForge produces tasks that are just the right difficulty to help agents learn better. Experiments show that training agents on these tasks improves their performance significantly compared to tasks designed by humans or simpler validation methods.
terminal agenttask synthesissolver calibrationmulti-solver calibrationcontrastive calibrationlearnabilityagent training dataterminal-benchexecutable validationadversarial training
Authors
Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
Abstract
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.