Flexiworld improves long-horizon control with flexible action chunks
FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
Machine Learning
Summary
Planning actions far into the future is challenging because decisions depend on many small steps. The authors present FlexiWorld, a system that learns and plans by grouping actions into chunks of varying lengths. This lets it better handle different time spans and goals at once. Their system improves success rates on several tests, compared to earlier methods with fixed-length action chunks. It also adapts its planning speed by using longer or shorter action chunks without needing retraining.
What this means in practice
- •For robotics engineers: Improve robot navigation and task planning by using flexible action chunks for more efficient long-term control over complex movements.
- •For autonomous vehicle developers: Enable more adaptable and robust decision-making systems by planning across multiple time scales with variable-length action sequences.
Authors
Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu
Abstract
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.