Alignment guided transformer improves robot vision language action tasks
Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning
Machine Learning
Summary
Robots that use vision, language, and actions to perform tasks often struggle because these three parts don’t always connect well. The authors introduce a new method called Alignment-Guided Flow Transformer (AGFT) that helps these three parts line up better, making the robot’s actions more accurate and easier to adapt to new tasks. Their approach also speeds up the robot’s decision-making without losing quality. Experiments show that this alignment method improves both success rates and efficiency in robot control.
What this means in practice
- •For robotics engineers: Train robot control systems that better integrate vision and language inputs for more accurate task execution and efficient adaptation.
- •For autonomous system developers: Design faster inference policy models for robots that reduce latency without sacrificing action accuracy.
Authors
Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao
Abstract
Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.