FluxVLA Engine simplifies building robots that see and act

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

RoboticsArtificial Intelligence

Summary

Building robots that understand what they see and act accordingly involves many complex parts like different data types, training systems, and ways to test. The authors created the FluxVLA Engine, which is a platform that connects all these pieces in one place to make it easier to develop and use robot policies. Instead of making a new robot model, FluxVLA standardizes how all components work together, from training to running on real robots. This helps turn new robot learning algorithms into practical, reliable systems faster.

What this means in practice

  • For robot software engineers: Build and deploy robot control systems that integrate visual, language, and action components with standardized interfaces to speed up development and testing.
  • For simulation platform developers: Use FluxVLA's dual-arm compositional simulation and automatic data generation to create more realistic and scalable robot testing environments.

Authors

Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen

Abstract

Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA