Robotic ultrasound navigation predicts anatomy changes with scene graphs
Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation
RoboticsArtificial Intelligence
Summary
Ultrasound imaging depends heavily on how well the operator can predict the anatomy seen as the probe moves. The authors created SonoGraph-WM, a system that uses a special map of body structures called scene graphs to predict future views and probe positions without generating images. It uses past data and a planning method to imagine and follow the best path to a target internal structure. They trained and tested their system using CT scans, showing it can reliably navigate robotically to organs like the gallbladder and pancreas. This approach could help make robotic ultrasound scanning more precise and less dependent on manual skill.
What this means in practice
- •For medical device engineers: Integrate anatomical scene-graph world models to improve robotic ultrasound probe path planning toward organ targets.$Commercial implications: Enables development of commercially deployable robotic ultrasound systems offering more reliable and autonomous scanning capabilities.
- •For medical simulation developers: Use CT-derived anatomical scene graphs to create training scenarios for robotic ultrasound navigation without extensive real ultrasound data.
Authors
Xuesong Li, Shuai Chen, Feng Li, Zhongliang Jiang, Nassir Navab, Yuan Bi
Abstract
Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/