Vision language action models improve robot motion understanding with frequencies
Frequency-Conditioned Flow Matching for Vision-Language-Action Models
Robotics
Summary
Robot actions happen over time and include moves at different speeds and sizes, which can be seen as having different 'frequencies.' Usually, computer models guess robot motions without paying much attention to these different frequencies. The authors created a new method called FreqFM that explicitly uses these frequency details to better predict actions. This approach fits into existing models and helps robots perform tasks more accurately in tests and real robots.
What this means in practice
- •For robotics engineers: Improve robot motion planning by using frequency-based modeling for more accurate action prediction.
- •For industrial automation teams: Enhance robotic task execution reliability in factories by integrating frequency-conditioned action models.
Authors
Haochen Niu, Shengye Dong, Hao Liu, Peiwen Lin, Wang Chuang
Abstract
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.