NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

2026-08-17Robotics

RoboticsArtificial Intelligence
AI summary

The authors address challenges in making Vision-Language-Action (VLA) models work efficiently and smoothly on different robots. They created NebulaVLA, which separates understanding what to do from how to do it, saving computing power and improving flexibility. To help different robots understand commands the same way, they developed a common action language called GESTURE-7. They also introduced a method to make robot movements smoother. Tests show that their approach works better and faster than previous ones.

Vision-Language-Action modelsasynchronous architecturesemantic reasoningaction controlcross-embodiment generalizationkinematic continuitymask-based smoothnessrobotic controlLIBERO-Plus dataset
Authors
Cong Zhao, Shuai Tian, Xu Zhang, Baocheng Ni, Xinguo Song, Xueying Sun, Shu Jiang, Shouchang Yang, Bo Tang, Jin Deng, Ge Zhu, YongCheng Wang, Jin Xu, Ri Yang
Abstract
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity. To bridge semantic gaps across heterogeneous robots, we introduce GESTURE-7, a unified language-grounded action representation. Furthermore, our Guide Action algorithm enforces kinematic continuity via mask-based smoothness constraints. Comprehensive evaluations demonstrate that NebulaVLA significantly outperforms synchronous baselines, achieving an 85.5\% average success rate on LIBERO-Plus and accelerating action generation by \textasciitilde 2.7$\times$. This asynchronous design enables highly efficient and responsive control for practical robotics.