Papers for

medical simulation developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SurgGMF forecasts future surgical scenes using Gaussian motion fields

SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering

Abstract: Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
Predicting what will happen next in a surgical scene can help robots assist better during operations. The authors created SurgGMF, a method that predicts future movement in surgery videos using a special way to represent motion called Gaussian motion fields. Instead of guessing future video frames directly, SurgGMF forecasts position, size, and rotation changes to show how things will move and appear next. Their method works better than traditional motion prediction techniques, allowing for more accurate and efficient anticipation of surgical scenes.
Open → 2609.34733v1

Robotic ultrasound navigation predicts anatomy changes with scene graphs

Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation

Abstract: Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/

Sat 26 SeptRoboticsArtificial Intelligence
The gist
Ultrasound imaging depends heavily on how well the operator can predict the anatomy seen as the probe moves. The authors created SonoGraph-WM, a system that uses a special map of body structures called scene graphs to predict future views and probe positions without generating images. It uses past data and a planning method to imagine and follow the best path to a target internal structure. They trained and tested their system using CT scans, showing it can reliably navigate robotically to organs like the gallbladder and pancreas. This approach could help make robotic ultrasound scanning more precise and less dependent on manual skill.
Open → 2609.32837v1

Video generation models struggle to simulate medical procedures accurately

REMEDY: How Far Is Video Generation from Medical Education World Models?

Abstract: Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current generators can meet these requirements has not been measured. To address this problem, we introduce Readiness Evaluation of Medical Education Demonstration sYnthesis (REMEDY), to our knowledge, the first benchmark for AI-generated medical teaching demonstrations. REMEDY provides 900 first frames from real demonstration videos, covering 12 tasks across four scenarios: operating room, imaging, clinic and bedside, and resuscitation. Five contemporary open-source video generation models produce 4,500 videos from these frames. We combine task-specific clinical checklists with video and motion quality metrics. Evaluation covers four dimensions: clinical action following, clinical profiles, video quality, and motion quality. Our results show that realistic appearance and temporal consistency do not ensure correct clinical actions. Even the most advanced MiniMax-H3 achieves only 28.25% on strict clinical success rate, and fine-grained clinical actions remain challenging. These findings establish a foundation and roadmap for developing future medical education world models.

Sat 26 SeptComputer Vision and Pattern Recognition
The gist
Generating videos that show medical procedures convincingly is very hard because the videos must look real and also show the correct steps. The authors created a new test called REMEDY to check how well AI models can do this. They tested five different video generators on 12 medical tasks but found even the best model got less than 30% of the clinical steps right. This shows current technology is not yet good enough for producing reliable medical teaching videos.
Open → 2609.32460v1

Probabilistic mesh estimation improves deformable object tracking

Differentiable Mesh State Estimation via Factor Graph Inference for Deformable Object Reconstruction

Abstract: Estimating deformable object states remains a fundamental challenge in robotics and simulation. We propose a novel factor graph-based framework for probabilistic mesh state estimation of deformable objects. The method directly updates a tetrahedral mesh, a rich and physically-grounded representation of an environment, by combining physics priors, noisy sensor measurements, and temporal smoothness constraints within a unified probabilistic formulation. The estimation problem is posed as a nonlinear least-squares optimization and solved using Levenberg-Marquardt. Ex vivo central-airway obstruction experiments and simulations on deforming cube models demonstrate reliable and accurate reconstruction under both rigid motion and deformation, highlighting the potential of this probabilistic approach for principled, measurement-driven mesh state estimation in deformable object reconstruction.

Tue 15 SeptRoboticsComputer Vision and Pattern Recognition
The gist
Tracking and understanding objects that change shape, like soft or bendable items, is hard for robots and simulations. The authors created a new method that uses a detailed 3D mesh made of tiny tetrahedrons to represent these objects. Their approach combines knowledge about physics, noisy sensor data, and smooth movement over time into a single calculation to figure out the object’s shape accurately. Tests on airway experiments and cube models show their method works well for objects that both move rigidly and deform. This could help robots better see and use soft or flexible things.
Open → 2609.16686v1