Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery
2026-08-03 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial Intelligence
AI summaryⓘ
The authors developed a system called DiffeoAfford to help surgeons by automatically identifying important areas to focus on during laparoscopic surgery. Their method uses computer tracking of tissue and surgical instruments from past surgeries to create labels showing where attention should be placed, without needing people to label each video frame manually. They trained a model with these labels that can predict important regions in real time and supports a tool called AffordView, which helps frame the surgical view automatically. Their approach matches expert opinions and surgeon eye movements and helps reduce the mental effort surgeons experience.
computational attention modelslaparoscopytissue trackinginstrument trajectoryaffordancediffeomorphismvisual attentionauto-framingcognitive workload
Authors
Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding
Abstract
Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical intent is highly specialized and tacit. Here, we introduce DiffeoAfford, an action-grounded tissue affordance framework that retrospectively derives visual attention supervision from completed surgical procedures. By combining diffeomorphism-constrained tissue tracking with instrument trajectory analysis, DiffeoAfford generates affordance hotspot labels without manual per-frame annotation. A real-time prediction model trained on these labels anticipates relevant surgical regions and enables AffordView, an assistive auto-framing system for laparoscopic visualization. The proposed framework aligns with expert annotations and intraoperative surgeon gaze, and reduces surgeon cognitive workload during real-world evaluations using subjective, physiological, and behavioral measures.