Papers for

medical device companies

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

SurgFlow improves surgical robot tool targeting using 3D object motion

SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation

Abstract: Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.

Sun 27 SeptRobotics
The gist
Surgical robots need to know not just how objects move but also when and where to touch them during surgery, which is a tough problem. The authors created SurgFlow, a system that learns from stereo surgery videos to predict both 3D object movements and when surgical tools should contact those objects without needing labeled actions. SurgFlow was tested on a common surgical robot, where it succeeded in almost all tasks, and it also worked well when transferred to a different robot and camera views. This approach helps surgical robots better understand and interact with tissues during procedures using widely available video data.
Open → 2609.33237v1

Foundation models show mixed reliability on retinal image tasks

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

Abstract: Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards

Sun 27 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Retinal image analysis helps detect eye diseases but AI models often struggle when tested on different patient groups or types of images. The authors created a benchmark called FOCUS that tests many AI models across multiple eye disease datasets from different places and imaging conditions. Their results show no single type of model works best everywhere; some models perform well generally, while others need fine-tuning to improve. This work helps to better understand how reliable and fair AI models are for eye disease diagnosis in diverse real-world settings.
Open → 2609.33158v1