CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

2026-08-31Robotics

Robotics
AI summary

The authors created CometVLA, a new model designed to improve robots' ability to understand and perform physical tasks by better linking vision, language, and action. They built CometData and CometBench, datasets that closely match how robots actually move and interact with the world, unlike previous separate and misaligned data. They also introduced a technique called Global Action Prior (GAP) tokens to help the model learn common motion patterns without messing up its understanding of physical concepts. Testing showed that their model performs better on real robot tasks and simulations, and that improving vision-language understanding helps robots act more successfully.

Vision-Language-Action (VLA) modelsPhysical CommonsenseEmbodied DataVisual Question Answering (VQA)Egocentric VideosGlobal Action Prior (GAP) tokensTeleoperationSimulationRobot ManipulationPre-training
Authors
Hanwen Wan, Dafeng Chi, Linbo Zhai, Tianao Shen, Yuzheng Zhuang, Tianle Zhang, Peidong Liu, Liang Lin, Xiaoqiang Ji
Abstract
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap. We construct CometData and CometBench, an embodied physical VQA corpus and benchmark strictly aligned with the robot's action data and embodiment. We introduce Global Action Prior (GAP) tokens, a compact learnable bottleneck that isolates task-agnostic motion regularities and lets the action head consume physical commonsense without corrupting the pre-trained VLM backbone. We co-train CometVLA across the embodied data pyramid, spanning teleoperation, simulation, egocentric trajectories, and VQA layers. On real-world manipulation tasks and RoboTwin simulation, CometVLA consistently outperforms strong VLA baselines. Correlation analysis shows that stronger VLM performance on CometBench indicates higher VLA success rates. Results demonstrate that physical understanding pre-training genuinely benefits downstream manipulation.