Large language models improve reasoning by optimizing skills prompts and routing
Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs
Artificial Intelligence
Summary
Language models can solve many tasks, but how you tell them what to do and which parts of their knowledge to use is very important. The researchers introduce a method called SPARO that improves language model programs by adjusting task instructions, reusable skills, and how the model chooses which skills to use. This method figures out which part to change when the model makes mistakes and updates that part to do better next time. Their tests show this approach works better than just changing the instructions or skills alone.
What this means in practice
- •For ai application developers: Improve AI reasoning systems by jointly refining prompts, skills, and usage policies to reduce errors and boost performance on complex tasks.
- •For automated customer support teams: Enhance modular chatbots to decide when and how to deploy specialized response skills, making troubleshooting more accurate and efficient.
Authors
Haoran Shou, Haoyue Liu, Yu Huo, Kun Zeng, Xiaoying Tang
Abstract
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either optimize a monolithic prompt, or separately induce and refine skills from model traces. As a result, they lack a principled mechanism for deciding which component should be updated when failures occur, and they rarely optimize prompts, skills, and skill-use policies in a unified framework. We propose SPARO (Skill, Prompt, And Routing Optimization), a framework that jointly optimizes task instructions, reusable skill blocks, and routing rules. It performs controlled counterfactual evaluations, converts examples' effects into a probabilistic responsibility distribution over prompt, skill, and routing components, samples one component from that distribution, and applies the corresponding targeted mutation. This design moves language-program optimization beyond global prompt rewriting toward structured, reusable, and selectively activated task knowledge. Across five benchmarks and five worker models, SPARO consistently outperforms both prompt-centered and skill-centered optimization baselines. These results suggest that effective language-program optimization depends not only on discovering useful task knowledge, but also on deciding where that knowledge should be stored and when it should be activated.