Decoupled vision language and action improve robot manipulation efficiency
Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies
Robotics
Summary
Many robot control systems use very large models that handle vision, language, and actions altogether, which makes them slow and energy-heavy. The authors show that separating or 'decoupling' the parts for seeing, understanding language, and deciding actions can work just as well but much faster and with less energy use. They tested this on simulated and real robots performing various tasks, finding comparable success to big combined models but with eight to seventeen times faster action steps. This approach makes robot control more efficient while still understanding language and vision well.
What this means in practice
- •For robotics engineers: Build real-time robot controllers that process vision and language separately for faster and less power-intensive manipulation tasks.
- •For industrial automation teams: Deploy more efficient robotic arms for diverse manipulation jobs with quicker response using decoupled vision-language-action policies.
- •For smart home device developers: Integrate efficient multi-modal robot control for home robots that understand voice commands and perform tasks quickly with lower energy use.$Commercial implications: Enables consumer home robots with responsive and energy-efficient multi-task abilities that improve user interaction and battery life.
Authors
Xiatao Sun, Chen Liang, Ziyao Zeng, Qian Wang, Haoyang Zhang, Yue Sun, Qiucheng Li, Daniel Rakita
Abstract
Vision-Language-Action (VLA) models attach an action module to a Vision-Language Model (VLM) with billions of parameters and pay for that backbone at every control step. For a low-level manipulation policy, this cost may be unnecessary: the VLM supplies vision and language embeddings, and recent standalone vision encoders and encoder-only language models now match or exceed large VLMs on visual embedding and language understanding benchmarks. We study this question with a controlled experiment. Holding the demonstrations, the training budget, the tasks, and the measurement platform fixed, we vary the vision encoder, the language encoder, and the action head of a decoupled policy and compare against seven VLA baselines. The study yields the Decoupled Embodiment Model (DEM), which pairs a fine-tuned DINOv3 encoder and a frozen NeoBERT encoder with a MeanFlow head that generates each action chunk in a single forward pass. On 18 simulated manipulation tasks with held-out language paraphrases and randomized scenes, and on three real-robot tasks, DEM achieves observed success comparable to state-of-the-art VLM-backbone policies under our evaluation protocol, while running at eight to seventeen times their inference frequency and drawing six to fifteen times less energy per inference. Within this task scope, modern decoupled components offer a better success--latency--energy trade-off.