Autonomous driving improves by aligning decisions with future scene views

RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

Robotics

Summary

Autonomous cars need to plan their driving by understanding what will happen next on the road. Previous methods tried to predict future scenes explicitly, which slowed learning and made driving decisions slower. The authors present RAF-VLA, a way to align the internal thinking of the car with future road views without actually predicting those scenes. This approach helps the car learn faster and drive just as well, while being quicker in making decisions. Tests show it matches top driving systems but with less training and only a tiny increase in computing time.

What this means in practice

  • For autonomous vehicle engineers: Improve autonomous car decision-making speed and training efficiency using internal future-view alignment without explicit future scene generation.
  • For robotics control developers: Develop faster control policies for robots that must plan based on visual input by aligning internal representations with expected future observations.

Authors

Dogun Kim, Yongjae Lee, Joonhee Lim, Yeina Lee, Junhyeok Park, Moogeun Park, Dongsuk Kum

Abstract

Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.