Vision language inputs show phase dependent effects on robot tasks
IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies
RoboticsArtificial Intelligence
Summary
Robots often use images, their own movements, and spoken instructions to complete tasks, but it can be unclear when and how each kind of input helps. The authors created a new way to trace how information from these different inputs influences a robot's actions at different stages of a task. Their method runs scenarios that change inputs slightly to see how the robot’s behavior and success are affected over time. This helps show when visual, language, or movement data are most important, and how they interact during the robot’s work.
What this means in practice
- •For robotics engineers: Identify which sensory inputs matter most at each step to improve robot task planning and reliability.
- •For autonomous vehicle developers: Analyze how different sensory data contribute over time to driving decisions for safer navigation.
Authors
Jinwoong Kim, Sangjin Park
Abstract
Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, and language instructions. However, it remains unclear at which execution stages each modality contributes to final task success and how input interventions propagate through subsequent states, observations, and actions. Existing attribution approaches primarily measure local sensitivity or temporally aggregated importance, limiting their ability to capture phase-dependent contributions and cross-phase dependencies. We propose Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies (IMPACT-VLA). IMPACT-VLA constructs behavioral phases from action transitions in a successful reference rollout, aligns them with policy query boundaries, and defines phase-modality blocks as attribution units. It then performs closed-loop counterfactual re-execution to quantify each block's contribution to final task success. We further analyze cross-phase non-additive interactions and trajectory propagation while distinguishing behavioral from functional recovery. Across 30 LIBERO robot manipulation tasks using OpenVLA-OFT, dominant-modality transitions occurred in 25 tasks (83.3%), and closed-loop attribution identified task-critical information more faithfully than Static Action Perturbation. Later-block marginal gains for negatively interacting pairs increased by approximately 3.3x under early-phase input replacement, while functional recovery could occur without behavioral recovery. These results reveal when multimodal inputs support task success and how their contributions become conditionally coupled during closed-loop execution.