Vision language models tested on coastal and underwater scenes

Evaluation of Vision-Language Models Across Diverse Coastal Environments

Computer Vision and Pattern RecognitionRobotics

Summary

Vision-language models help robots understand pictures by linking images to words. This paper studies how well these models work with photos from coastal areas in Hawaii, which can be different from land scenes. The authors made a special dataset of over 1,000 coastal images with detailed labels and tested seven popular models. They found that models recognize broad landscapes better than specific coastal objects, and special labels can help improve recognizing some coastal things. The challenges come mainly from how objects are segmented in images and described by words.

What this means in practice

  • For robotics engineers: Use coastal-specific datasets and label improvements to better train perception systems for robots operating in shoreline environments.
  • For environmental monitoring teams: Improve automated identification of coastal features in images to assist in ecosystem monitoring and management tasks.

Authors

Seth Knoop, Chad R. Samuelson, Gabriel R. Slade, Brady Moon, Joshua G. Mangelson

Abstract

Vision-language models (VLMs) enable robotic per- ception by associating visual observations with natural-language concepts. Yet their performance in coastal environments remains largely unexplored. We introduce a densely labeled coastal dataset containing more than 1,000 images collected across seven missions in three regions of Oahu, Hawaii, with 18 semantic classes and over 7,400 annotated instances. We evaluate seven modern VLMs through three complementary experiments mea- suring text-to-mask, mask-to-mask, and mask-to-text alignment. Broad landscape classes are generally recognized more accurately than conventional object and coastal classes, with coastal con- cepts presenting the greatest challenge. However, comparisons of shared conventional classes across coastal and terrestrial datasets reveal no consistent performance difference attributable solely to environmental context. Mask-to-mask matching also remains similar across conventional and coastal classes, while alternative textual labels substantially improve recognition of several coastal concepts. These results suggest that lower performance on coastal classes (at least on the objects/query categories evaluated) is heavily influenced by segmentation and linguistic representation.