Papers for
medical simulation developers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
SurgGMF forecasts future surgical scenes using Gaussian motion fields
SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering
Abstract: Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.
Robotic ultrasound navigation predicts anatomy changes with scene graphs
Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation
Abstract: Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/
Video generation models struggle to simulate medical procedures accurately
REMEDY: How Far Is Video Generation from Medical Education World Models?
Abstract: Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet these requirements has not been measured. To address this problem, we introduce Readiness Evaluation of Medical Education Demonstration sYnthesis (REMEDY), to our knowledge, the first benchmark for AI-generated medical teaching demonstrations. REMEDY provides 900 first frames from real demonstration videos, covering 12 tasks across four scenarios: operating room, imaging, clinic and bedside, and resuscitation. Five contemporary open-source video generation models produce 4,500 videos from these frames. We combine task-specific clinical checklists with video and motion quality metrics. Evaluation covers four dimensions: clinical action following, clinical profiles, video quality, and motion quality. Our results show that realistic appearance and temporal consistency do not ensure correct clinical actions. Even the most advanced MiniMax-H3 achieves only 28.25% on strict clinical success rate, and fine-grained clinical actions remain challenging. These findings establish a foundation and roadmap for developing future medical education world models.
Probabilistic mesh estimation improves deformable object tracking
Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction
Abstract: Estimating deformable object states remains a fundamental challenge in robotics and simulation. We propose a novel factor graph-based framework for probabilistic mesh state estimation of deformable objects. The method directly updates a tetrahedral mesh, a rich and physically-grounded representation of an environment, by combining physics priors, noisy sensor measurements, and temporal smoothness constraints within a unified probabilistic formulation. The estimation problem is posed as a nonlinear least-squares optimization and solved using Levenberg-Marquardt. Ex vivo central-airway obstruction experiments and simulations on deforming cube models demonstrate reliable and accurate reconstruction under both rigid motion and deformation, highlighting the potential of this probabilistic approach for principled, measurement-driven mesh state estimation in deformable object reconstruction.