Vision language action models shrink drastically with fast offline recovery

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

Robotics

Summary

Robots use big AI models that understand vision and language to follow instructions, but these models are often too large to run on actual robots efficiently. The authors show how to shrink these models by removing many parts and then quickly restore most of their capability using a method that only needs offline data. Their approach narrows the model’s size without losing much success in tasks, running faster and using less memory than before. This makes it easier to put powerful robot AI on real hardware without lengthy retraining.

What this means in practice

Authors

Chiyoung Kim, Sanghyuk Roy Choi, Minhyeok Lee

Abstract

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%. A recent approach restores such a model with supervised fine-tuning followed by reinforcement learning, which needs online rollouts and hundreds of GPU-hours. We recover most of the lost success entirely offline. Width pruning narrows the blocks but keeps the residual stream at its original size, so teacher and student hidden states have the same shape and are matched directly, without a projector. Training against a cache built in one teacher pass lifts the 63%-reduced student to within 3.5 points of the teacher in about 8 GPU-hours. A sweep over nine ratios locates where the recovery objective starts to matter. Up to 45% reduction the two do not differ significantly on OpenVLA-OFT. Hidden-state distillation then adds +2.1 to +4.5 points there between 63% and 87%, and +9.4 to +22.1 points on CogACT from 63% onward. At 81% on CogACT, a tripled recovery budget narrows the distilled student's gap to the teacher to 3.9 points on average, while supervised recovery stays more than 20 points below. At matched compression, width pruning yields higher success and depth pruning lower latency. On a 6-DoF manipulator, the distilled student at 72% reduction reaches 77.5% success against 59.5% for supervised recovery, runs 2.23x faster on-board than the teacher, and uses 62% less memory.