Efficient vision language control improves robot tasks with token caching
Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference
Computer Vision and Pattern RecognitionRobotics
Summary
Robots that understand both pictures and words often need a lot of computing power to decide what to do. The authors found a way to make this faster without retraining the robot’s brain by saving and reusing certain important visual and text clues. Their method, called TVCache, picks the most helpful parts of the robot’s thought process to speed up decision-making, especially by focusing on how words and images work together. Tests showed robots using TVCache completed tasks more successfully and with less computing work.
What this means in practice
- •For robotic engineers: Improve real-time robot control by reducing computation while maintaining task success rates through smart token caching techniques.
- •For autonomous vehicle developers: Speed up vision-language models interpreting sensor and map data by reusing stabilized representations to enable faster action planning.
Authors
Qianer Li, Chengjie Zhang, Jingwen Chen, Zanjia Tong, Jiyuan Zhang, Hong Zhang
Abstract
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.