Driving model adapts reasoning depth and focus by scene risk

CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

Computer Vision and Pattern Recognition

Summary

Driving decision systems often decide if they should think more about a situation but ignore how their thinking should change depending on what’s around and how risky it is. The authors present CAR-VLA, a driving model that decides not only how long to think but also how to think based on scene complexity and risk. It switches between quick responses, careful planning, or focused emergency reactions depending on how dangerous and complicated the driving scenario is. They trained and tested CAR-VLA in various simulated environments, showing it can drive well and respond appropriately to hazards.

What this means in practice

  • For autonomous vehicle developers: Create self-driving systems that adjust their decision-making style based on complexity and risk to improve safety and efficiency.$Commercial implications: Enables adaptable driving AI that meets safety and performance needs, supporting commercial autonomous vehicle products.
  • For robotics engineers: Integrate risk-aware reasoning modes into robots navigating dynamic environments for improved real-time hazard response.

Authors

Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Shi Fan, Mengyang Zhao, Wenjuan Meng, Zisheng Chen, Zhihao Zhu, Zhounan Jin, Hengli Wang, Qingfan Wang, Jiamei Liang, Bin Li, Xiangyang Xue

Abstract

Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git