Multimodal agents improve decision making with co-evolved skill guidance
SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
Artificial Intelligence
Summary
Multimodal agents, which use different types of input like images and text, need clear guidance to make good decisions over many steps. The authors noticed that a well-made skill can both guide how the agent acts and clearly define what success looks like. They created SkillRubric, which pairs directions for the agent with a way to evaluate progress using screenshots and outputs. By improving guidance and evaluation together in cycles, their method helped agents perform better in tests involving planning and tool use.
What this means in practice
- •For multimodal ai developers: Improve interactive agents' performance by integrating co-evolved skill instructions with evaluators for better planning and tool use.
- •For automation engineers: Enhance automated systems by refining step-by-step process guidance alongside progress evaluation to handle complex workflows more effectively.
Authors
Bingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang, Bingning Wang, Tianyi Lin, Zichao Yu, Yujin Han, Ziye Ma, Difan Zou
Abstract
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervision for intermediate decisions. Rubric-based rewards address this limitation through explicit intermediate criteria, but reliable rubrics are difficult to construct at scale and often disconnected from the procedure followed by the actor. We observe that a well-structured skill naturally specifies both how to act and what successful execution should achieve. Based on this insight, we introduce SkillRubric, which represents each skill through aligned actor-facing guidance and an evaluator-facing rubric. A multimodal verifier evaluates skill-defined goals using screenshots and tool outputs, assigning completion and progress rewards to the responsible turns. We further introduce an alternating co-evolution scheme that validates guidance revisions through paired rollouts under a frozen policy and rubric revisions offline under fixed guidance. Experiments across diverse multimodal agent benchmarks demonstrate consistent performance gains, while controlled paired rollouts further show that evolved skills provide more effective guidance for planning and tool use than their preceding versions.