Multi agent workflows learn when to skip steps for efficiency
Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
Artificial Intelligence
Summary
Large language model workflows often use several steps like planning, checking, and summarizing to get better answers. But running all these steps every time can waste work or even mess up good results. The authors created a method called Learning What to Skip (LW2S) that figures out when some steps can be safely left out by learning from past runs. This helps save computing resources while keeping or improving accuracy across tasks like math problems and coding.
What this means in practice
- •For software developers: Reduce computation time for multi-step AI tasks by skipping unnecessary workflow components safely.
- •For computational linguists: Optimize multi-agent language model pipelines by learning which language model components can be omitted without losing accuracy.
Authors
Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang, Jinze Li, Zheng Xing, Mengran Li, Victor C. M. Leung
Abstract
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.