FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation

2026-07-09Robotics

Robotics
AI summary

The authors introduce FabriVLA, a lightweight model that helps robots perform many different tasks by understanding vision and language together and then deciding on actions. It uses a special combination of a vision-language model and an action component that pays attention to important details. They trained it all at once starting from a model that already understands images and words, and tested it on 50 different tasks where it did very well. Their results show that you don't need extremely large models to get strong robot performance on many tasks.

Vision-Language Model (VLM)Action HeadFlow-MatchingSelf-AttentionMulti-Task ManipulationMeta-World MT50Joint OptimizationInternVL3.5Spatial ContextRobotics
Authors
Shiyuan Yang, Borong Zhang, Jizheng Zhang, Zhijia Tao, Junfei Guo, Donglai Ran, Xu Bian, Qingbiao Li
Abstract
We present FabriVLA, a lightweight Vision-Language-Action model for Precise Multi-Task Manipulation. FabriVLA combines an InternVL3.5 vision-language backbone with a flow-matching action head featuring gated self-attention across action tokens and shallow VLM layer fusion for enriched spatial context. The model is trained via single stage joint optimization from a pretrained VLM and randomly initialized action head. On the Meta-World MT50 benchmark spanning 50 diverse manipulation tasks, FabriVLA achieves a tier-average success rate of 90.0%, demonstrating that a compact VLA built on a 1B scale VLM can achieve strong performance without relying on multi billion parameter VLA backbones.