MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors present a new method to separate objects in videos using sounds without needing to train any new models. They combine large language models that understand multiple types of information with special segmentation models that outline objects. Their approach breaks the task into steps and uses existing models to link sounds and visuals effectively. This method works well on a video segmentation challenge focused on audio cues.

audio-guided video segmentationMultimodal Large Language ModelsSAM-based segmentation modelsfoundation modelstext-visual correspondenceobject mask generationtraining-free frameworkMeViS-Audio TrackLSVOS Challenge
Authors
Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu
Abstract
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.