Large diverse surgical video dataset improves instrument segmentation
LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation
Computer Vision and Pattern Recognition
Summary
Surgical videos often include lots of instruments, and doctors or computers need to identify these tools from descriptions. The authors found that existing datasets are small and limited, only referring to one instrument at a time. They created a much bigger dataset with many videos, instrument types, and different ways to describe the instruments, making it easier to train and test better computer models. They also tested several existing models and introduced a new method that showed promising results.
What this means in practice
- •For medical imaging developers: Build more robust surgical instrument identification systems by training on a large, diverse, and richly annotated video dataset.
- •For surgical robotics teams: Improve detection and tracking of multiple surgical tools simultaneously during procedures using textual references.
Authors
Zan Wang, Yunhe Feng, Dong Nie, Oluwatosin Oluwadare, Kewei Sha, Yan Huang, Heng Fan
Abstract
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.