Self supervised video tracking improves surgery without annotations
S3-Tracker: Self-Supervised Surgical Tissue Tracking With Contrastive Random Walks
Computer Vision and Pattern RecognitionMachine LearningRobotics
Summary
Tracking points on soft and moving tissue during surgery is very important but hard to teach computers because it's difficult to label videos with exact point movements. The authors created a method that learns to track points in surgical videos without needing labeled examples by figuring out pixel matches across frames on its own. This method performs about as well as approaches that use some supervision, and it naturally handles tissue changes during surgery. Their work shows it's possible to track surgical tissue points accurately without expensive annotations.
What this means in practice
- •For surgical robotics teams: Integrate annotation-free point tracking into surgical robots to maintain accurate video-to-imaging alignment despite tissue movement.
- •For medical device developers: Develop navigation systems that track soft tissues continuously in endoscopic surgery videos without manual labeling of training data.
- •For medical imaging software vendors: Create enhanced surgical support software that improves tissue tracking performance with reduced annotation costs.$Commercial implications: Enables selling software upgrades that improve tracking accuracy while lowering annotation expenses for hospitals.
Authors
Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean
Abstract
Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous registration between intraoperative video and preoperative imaging despite soft tissue deformation. However, supervised tracking methods depend on large annotated datasets, while surgical conditions make reliable trajectory annotation challenging. We propose a self-supervised Track-Any-Point approach that learns from unlabeled surgical videos by establishing global pixel correspondences and inferring point trajectories through contrastive random walks. Trained without annotations, our method achieves performance comparable to existing semi-supervised approaches while implicitly handling tissue deformation. These findings demonstrate the feasibility of self-supervised point tracking in surgical environments and its potential to reduce reliance on annotated data.