Agent adapts learning priorities to fix evolving skill gaps
ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Artificial Intelligence
Summary
As an AI agent learns over time, the problems it struggles with change, making fixed training goals less helpful. The authors propose a method called ARISE that constantly updates what the agent should focus on by using what it notices in its own behavior. ARISE develops flexible scoring rules and skills that evolve together to explore new challenges and improve weaker abilities. This approach helps the agent learn tasks better and faster in complex environments.
What this means in practice
- •For robotics engineers: Improve robot learning by continuously adjusting training goals to target newly emerging weaknesses in task performance.
- •For game ai developers: Enhance non-player character learning by dynamically evolving evaluation criteria and exploration strategies for complex gameplay.
Authors
Kun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu, Yongxiang Zhao, Yu Liu, Xingyu Lu, Lintao Ma, Kan Ren
Abstract
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .