SkillAA improves AI skill updating with precise graph-based editing
SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback
Artificial Intelligence
Summary
AI systems often use fixed skills to perform tasks, but fixing errors in those skills can be tricky and inefficient. The researchers created SkillAA, a method that uses a detailed graph to represent skills and how they connect, allowing the system to locate the exact part that caused a failure and update only that part. This makes AI skill updates more accurate and reliable. Testing with a powerful language model showed SkillAA achieves high success rates on several question-answering tasks.
What this means in practice
- •For ai application developers: Update AI task skills more precisely by targeting problematic graph parts for repair and validation in frozen language models.
- •For customer support automation teams: Improve automated question-answering systems’ correctness by selectively modifying procedural skills based on failure analysis.
Authors
Ziqiao Shang, Ling-Yue Ge, Lan-Zhe Guo
Abstract
External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.