Papers for

surgical robotics teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SurgGMF forecasts future surgical scenes using Gaussian motion fields

SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering

Abstract: Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Predicting what will happen next in a surgical scene can help robots assist better during operations. The authors created SurgGMF, a method that predicts future movement in surgery videos using a special way to represent motion called Gaussian motion fields. Instead of guessing future video frames directly, SurgGMF forecasts position, size, and rotation changes to show how things will move and appear next. Their method works better than traditional motion prediction techniques, allowing for more accurate and efficient anticipation of surgical scenes.
Open → 2609.34733v1

SurgFlow improves surgical robot tool targeting using 3D object motion

SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation

Abstract: Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.

Sun 27 SeptRobotics
The gist
Surgical robots need to know not just how objects move but also when and where to touch them during surgery, which is a tough problem. The authors created SurgFlow, a system that learns from stereo surgery videos to predict both 3D object movements and when surgical tools should contact those objects without needing labeled actions. SurgFlow was tested on a common surgical robot, where it succeeded in almost all tasks, and it also worked well when transferred to a different robot and camera views. This approach helps surgical robots better understand and interact with tissues during procedures using widely available video data.
Open → 2609.33237v1

Large diverse surgical video dataset improves instrument segmentation

LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation

Abstract: Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.

Sat 19 SeptComputer Vision and Pattern Recognition
The gist
Surgical videos often include lots of instruments, and doctors or computers need to identify these tools from descriptions. The authors found that existing datasets are small and limited, only referring to one instrument at a time. They created a much bigger dataset with many videos, instrument types, and different ways to describe the instruments, making it easier to train and test better computer models. They also tested several existing models and introduced a new method that showed promising results.
Open → 2609.23067v1

Self supervised video tracking improves surgery without annotations

S3-Tracker: Self-Supervised Surgical Tissue Tracking With Contrastive Random Walks

Abstract: Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous registration between intraoperative video and preoperative imaging despite soft tissue deformation. However, supervised tracking methods depend on large annotated datasets, while surgical conditions make reliable trajectory annotation challenging. We propose a self-supervised Track-Any-Point approach that learns from unlabeled surgical videos by establishing global pixel correspondences and inferring point trajectories through contrastive random walks. Trained without annotations, our method achieves performance comparable to existing semi-supervised approaches while implicitly handling tissue deformation. These findings demonstrate the feasibility of self-supervised point tracking in surgical environments and its potential to reduce reliance on annotated data.

Sun 13 SeptComputer Vision and Pattern RecognitionMachine LearningRobotics
The gist
Tracking points on soft and moving tissue during surgery is very important but hard to teach computers because it's difficult to label videos with exact point movements. The authors created a method that learns to track points in surgical videos without needing labeled examples by figuring out pixel matches across frames on its own. This method performs about as well as approaches that use some supervision, and it naturally handles tissue changes during surgery. Their work shows it's possible to track surgical tissue points accurately without expensive annotations.
Open → 2609.14313v1