SurgFlow improves surgical robot tool targeting using 3D object motion
SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation
Robotics
Summary
Surgical robots need to know not just how objects move but also when and where to touch them during surgery, which is a tough problem. The authors created SurgFlow, a system that learns from stereo surgery videos to predict both 3D object movements and when surgical tools should contact those objects without needing labeled actions. SurgFlow was tested on a common surgical robot, where it succeeded in almost all tasks, and it also worked well when transferred to a different robot and camera views. This approach helps surgical robots better understand and interact with tissues during procedures using widely available video data.
What this means in practice
- •For surgical robotics teams: Automate precise tool motions and contact timing in robot-assisted surgeries using only stereo video inputs without manual action labeling.
- •For medical device companies: Integrate SurgFlow-based contact flow prediction to improve the autonomy and reliability of robotic surgical tools across different platforms and camera setups.$Commercial implications: Enables development of autonomous surgical systems adaptable to multiple robots, enhancing product capabilities in robotic surgery.
Authors
Changwei Chen, Xiao Liang, Yinuo Yang, Nicole Shen, Peihan Zhang, Sara Wickenhiser, Zekai Liang, Soofiyan Atar, Michael Yip
Abstract
Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.