Robot policies reuse planned actions to reduce computation calls
Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
RoboticsArtificial IntelligenceComputer Vision and Pattern RecognitionMachine Learning
Summary
Robot programs often plan several future moves but only use the first few before making a new plan, which can be slow or unresponsive. The authors found that the unused planned moves can often still be useful and safe to perform, as long as movements don't suddenly change. They created a simple method that reuses these leftover moves without extra calculations or changing the robot’s software. This approach speeds up the robot's decision-making without lowering success in completing tasks.
What this means in practice
- •For robotics developers: Increase efficiency of robot control systems by reusing planned action sequences to reduce policy computation without sacrificing task success.
- •For manufacturing automation teams: Deploy faster robot manipulation in factory settings by minimizing control overhead using action upcycling on existing policy chunks.
Authors
Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye
Abstract
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose *Action Upcycling*, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2--1.7$\times$ with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.