MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

2026-07-27Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionArtificial Intelligence
AI summary

The authors created a new marine video dataset called MarineEVT to help computers better understand important events in underwater videos, which are hard to analyze because these events happen rarely and unpredictably. They developed a method called EVT-R1 that uses special visual tools to help the model focus on critical moments and answer questions about the videos. Their approach performed better than 11 other top models, showing promise for improving marine video analysis and education. This work supports understanding ocean life and ecological interactions using video technology.

Vision-Language ModelsMarine Video UnderstandingVisual Question AnsweringEvent-centric DatasetTemporal UnderstandingVisual ToolsEcological InteractionsVideo AnalysisDomain ExpertiseSustainable Ocean Monitoring
Authors
Tuan-An To, Yuk-Kwan Wong, Tuan-Anh Vu, Ziqiang Zheng, Sai-Kit Yeung
Abstract
Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.