Papers for

autonomous vehicle control teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

ActionUNet improves vision language action robot control success rates

ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning

Abstract: Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves π0.5 success rates by absolute 9.8%, 6.1%, and 11.4%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code and implementation details are available at https://github.com/Di-Zhu123/ActionUNet.

Mon 28 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Robots that understand instructions combining images and language often struggle to perform precise, smooth actions in messy environments. The authors propose ActionUNet, a method that fine-tunes existing robot models efficiently to better connect overall instructions with detailed movements. ActionUNet uses a special network design to combine different time scales and a smooth action decoder to reduce shaky motions. This helps robots perform tasks more reliably in tests and real-world conditions without losing their general ability.
Open → 2609.34982v1