Robotic foundation models struggle with visual challenges in manipulation tasks

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

RoboticsComputer Vision and Pattern Recognition

Summary

Robots that rely on visual information to perform tasks often assume they have perfect, up-to-date views of their surroundings. This paper presents a new test called LIBERO-VPro that introduces different kinds of visual problems to see how well these robot models can still perform. The authors found that while robots do well in ideal conditions, their performance drops sharply when the visuals are noisy, outdated, or inconsistent. Different robot models have varying weaknesses, showing that how robots process visual information matters for their reliability.

What this means in practice

  • For robotics engineers: Improve robot vision systems by identifying and addressing specific visual weaknesses during real-world object manipulation tasks.
  • For industrial automation teams: Design more reliable robotic manipulators that maintain performance despite camera delays, occlusions, or unexpected visual changes in factory settings.

Authors

Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, Bin Zhu

Abstract

Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.