Robot plans updated on the fly improve task success rates
Revision, Not Restart: Revisable Visual Plans for Closed-Loop World-Action Models
RoboticsComputer Vision and Pattern Recognition
Summary
Robots often make plans based on what they expect to see and do, but sometimes their real experiences don’t match those plans. The authors propose a way for robots to revise their visual plans when new information comes in, instead of starting over from scratch. This method helps robots adjust their actions more smoothly and improves how often they succeed at tasks. They tested this approach on robot benchmark tasks and saw noticeable improvements.
What this means in practice
- •For robotics engineers: Improve robots' ability to adapt plans mid-action based on real-time feedback, increasing task success in unstructured environments.
- •For industrial automation teams: Enhance robotic systems by integrating revisable visual plans that reduce the need to restart tasks after errors or unexpected events.
Authors
Pengyiang Liu, Junbo Niu, Wenhao Zheng, Xinchen Chen, Canyu Li, Zhongyue Shi, Jiahao Xie, Si Liu
Abstract
World-action models use predicted visual futures to condition robot actions, yet execution feedback can invalidate parts of a prediction while leaving its task structure useful. We propose Revisable Temporal Planning (RTP), which maintains the visual future as a persistent action condition and revises it after feedback. Its central mechanism is a learned revision bridge: it resumes an intermediate state saved during visual generation and adapts its continuation to current observations. Visual and action supervision connect this revision to subsequent control. Time-aware history supplies observed evidence, and an adaptive policy selects retention, bridge revision, or fresh replanning from new noise before decoding the next action. On RoboMME and RMBench, RTP achieves task-averaged success rates of 48.6% and 84.8%, respectively. Matched comparisons support learned continuation; estimated checkpoint-source and action-prefix effects are positive but less precisely resolved. These results connect feedback-driven visual-plan revision to closed-loop task performance. Project Page: https://PLACEHOLDER.github.io/RTP/