Mbariml tool helps label deep-sea images for better AI detection models

mbariml: a curation pipeline for turning deep-sea imagery and video into object-detection training data

Computer Vision and Pattern Recognition

Summary

Identifying animals in deep-sea photos and videos is hard because they are rare and look faint. The authors created mbariml, a tool that runs an AI detector on these images and videos, then groups detected objects by how they look so a person can quickly approve or correct them. This helps create better labeled data to train AI models to spot these deep-sea creatures more accurately. The system also manages video differently by choosing key frames to reduce redundant data for review.

What this means in practice

  • For marine biologists: Create labeled training data sets for detecting and studying sparse marine organisms in underwater videos and images.
  • For computer vision engineers: Improve the efficiency and accuracy of labeling object detection datasets from underwater video by grouping similar detections for fast human review.

Authors

Lonny Lundsten, Kevin Barnard, Dave Caress

Abstract

Training data quantity and quality greatly affect object detection model performance, regardless of model architecture. When using object detection models on video and images from the deep sea, in which the objects of interest, primarily organisms, are sparse, faint, and hard to identify, incremental improvements to object detector performance may require an iterative approach to data labeling and management. This paper presents mbariml, a python-based video and image analysis pipeline built around the data labeling management process. mbariml uses an Ultralytics YOLO detection model, runs it over still images or video, stores every detection as a reviewable region of interest, groups those regions by visual similarity so that a human can accept or reject them in bulk, and exports the result as training data, statistics, image sidecars, and additional metadata. The human review stage is the centre of the design: an annotator can validate, relabel, resize, delete, and draw entirely new localizations, and every one of those edits is written back to the same database the detector wrote to. Video receives particular attention: the software treats each tracker-produced track as a provisional observation and selects one representative frame instead of retaining every detection in the track. We describe the pipeline stage by stage, including the operational middle-third heuristic used for track observation selection.