SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment
2026-08-31 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors created SurgSkill-Bench, a new dataset combining videos of surgical training simulations with expert skill scores and written feedback. They tested how well computer models can predict surgical skill scores just from video, and also when expert comments are included. Using smart video frame selection improved predictions based on video alone, and adding expert comments helped even more. They also discuss some limits like dataset size and how to interpret results when comments are used.
surgical skill assessmentOSATS scoresvideo analysismachine learningexpert commentskey-frame samplingvideo-text fusionAUROC
Authors
Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang
Abstract
Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on labor-intensive expert review. Existing automated approaches primarily focus on visual inputs and provide limited support for jointly studying operative performance, structured skill scores, and evaluator feedback. We introduce SurgSkill-Bench, an initial video-score-text benchmark-style dataset containing 214 surgical training simulation videos, six-dimensional OSATS scores, and expert free-text comments. We define two evaluation settings: video-only OSATS prediction for automated assessment and post hoc expert-comment-assisted prediction, where evaluator comments are available as auxiliary information. We provide controlled baseline experiments using representative frozen visual backbones, content-adaptive key-frame sampling, and a simple video-text co-attention fusion module. Under internal video-level validation, content-adaptive sampling improves video-only performance in this dataset, while evaluator comments provide additional score-related signal in the assisted setting. The best mean AUROC reaches 0.88 under dataset-specific median dichotomization. We further discuss evaluation constraints related to dataset scale, metadata completeness, and the interpretation of comment-assisted prediction. Code will be released publicly at a later date.