Vista harness improves multimodal reasoning in visual environments

VISTA: A Visual Harness for Reasoning in an Interactive World

Artificial IntelligenceComputer Vision and Pattern Recognition

Summary

Some AI models can understand and reason about pictures, but they struggle with long sequences of events. The authors created VISTA, a tool that helps these models remember and organize what they see over time, allowing better decision-making. With VISTA, an AI model performed perfectly on a set of visual games and used fewer actions than humans did. This tool can also work well across different types of visual puzzles and games with little adjustment.

What this means in practice

  • For interactive game developers: Enable AI agents to better understand and solve complex visual game challenges by efficiently handling long sequences of visual information.
  • For robotics engineers: Improve robot perception and reasoning by integrating a memory system that preserves and retrieves visual data during tasks in changing environments.

Authors

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

Abstract

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.