Vision language action models can mask weak target sensitivity
Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
RoboticsComputer Vision and Pattern RecognitionMachine Learning
Summary
This paper shows that even when robot vision models seem accurate in predicting object states, they might not truly react to changes in the object's physical position. The authors designed tests that separate how well the model predicts from how sensitive it is to real changes and to context. They found models can do well on natural inputs but still not respond properly when those inputs change in controlled ways. Adding measures of how well models respond to changes improved predicting robot failures, suggesting that checking responsiveness gives better insight than accuracy alone.
What this means in practice
- •For robotics engineers: Improve robot control systems by evaluating state prediction responsiveness to prevent errors during tasks.
- •For quality assurance teams in automation: Use responsiveness and context sensitivity metrics to enhance failure prediction in automated vision-language systems.
Authors
Hyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee, Seungmin Rho
Abstract
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary together, allowing contextual cues to contribute to prediction. In this paper, we introduce an evaluation framework that separates prediction accuracy, target-state responsiveness, and context stability using physically validated observations that cross target coordinates with robot contexts. We demonstrate that high natural-trajectory accuracy can coexist with weak controlled target-state responsiveness in fixed representation-readout pairs. Comparisons and interventions involving representations, readouts, and training data show that the three properties provide distinct diagnostic information. Furthermore, adding responsiveness and context sensitivity to a failure predictor based on initial state error and physical variables reduces policy-failure prediction error on new initializations relative to the specified baseline while same-observation controlled MAE is also informative. These findings motivate evaluating target-state responsiveness and context stability alongside natural prediction accuracy, and examining their relationship to actual policy behavior and task outcomes.