ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
2026-08-03 • Robotics
Robotics
AI summaryⓘ
The authors studied how robots perform tasks that require remembering past actions while adjusting their movements. They found existing methods either remember long-term goals or short-term actions but don't connect these well. Their new system, ChainVLA, keeps a running memory of what has happened and plans upcoming actions based on both recent movements and task progress. This approach helped ChainVLA achieve much better success in manipulation tasks, showing both parts of their method are important. Removing either memory or action continuity caused a big drop in performance.
vision-language-action (VLA) policieslong-horizon manipulationrecurrent stateevent memoryaction horizonrobotic task planningRMBenchLIBERO suitemotion continuity
Authors
Yuzhi Huang, Weijue Bu, Ziyi Xiong, Jie Wu, Fanding Huang, Jingyan Jiang, Zhi Wang
Abstract
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term task evidence through memory or short-term motion through action reuse and ensembling, leaving the cross-query handoff incomplete. We introduce ChainVLA, a 1.2B-parameter VLA policy that chains successive queries through a joint and revisable execution state. Progress Context combines a recurrent Working State with sparse event memory to carry observation-derived task progress, while Motion Tail feeds the preceding prediction's unexecuted continuation into state construction and action generation. Together, the two components condition a decoder that regenerates each action horizon under the latest observation, allowing the carried state to guide the next prediction without fixing it. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites, while removing Motion Tail or Progress Context reduces RMBench success to 11.2% and 3.0%, respectively. These asymmetric ablations are consistent with motion continuity helping preserve the observation stream from which task progress is inferred.