SurgGMF forecasts future surgical scenes using Gaussian motion fields
SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering
Computer Vision and Pattern Recognition
Summary
Predicting what will happen next in a surgical scene can help robots assist better during operations. The authors created SurgGMF, a method that predicts future movement in surgery videos using a special way to represent motion called Gaussian motion fields. Instead of guessing future video frames directly, SurgGMF forecasts position, size, and rotation changes to show how things will move and appear next. Their method works better than traditional motion prediction techniques, allowing for more accurate and efficient anticipation of surgical scenes.
What this means in practice
- •For surgical robotics teams: Improve robotic systems with better anticipation of patient scenes during surgery to enhance decision support and simulation.
- •For medical simulation developers: Create more realistic and dynamically predictive surgical simulations by forecasting tissue and instrument motions in advance.
Authors
Jingqian Sun, Yichao Tang
Abstract
Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.