InfiMed2 improves medical image and text understanding with new training approach

InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision

Computation and LanguageComputer Vision and Pattern Recognition

Summary

Medical information often comes in pictures and text, but combining these to help doctors can be tricky. The authors created InfiMed2, a computer program that learns from a large collection of medical texts and images in a careful step-by-step way. They improved how it learns to answer medical questions clearly and accurately. InfiMed2 performed well on several medical tests, even beating some bigger models.

What this means in practice

  • For healthcare ai developers: Develop AI systems that interpret both medical images and texts for improved diagnostic support using InfiMed2’s stage-aware training and supervision techniques.
  • For clinical software engineers: Integrate advanced multimodal medical foundation models into clinical tools to provide more accurate and explainable answers from patient data.$Commercial implications: Enables new AI-powered diagnostic and decision-support software for hospitals and clinics with stronger medical reasoning capabilities.

Authors

Guanghao Zhu, Zeyu Liu, Zhitian Hou, Pengkai Wang, Zhijie Sang, Shuo Cai, Yang Yu, Yuanyi Wang, Yanggan Gu, Congkai Xie, Jianmin Wu, Hongxia Yang

Abstract

Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information density, and their utility shifts as training progresses from broad knowledge acquisition to late-stage consolidation. Meanwhile, post-training is often dominated by short-form visual question answering, providing limited supervision for informative and answer-consistent explanations. We introduce InfiMed2, a family of 4B and 27B generalist medical multimodal foundation models built around stage-aware data design. We curate a 55.68B-token corpus that combines broad clinical knowledge with context-rich biomedical visual evidence through source-specific processing. Our CPT pipeline first adapts the vision encoder, then builds broad medical knowledge, and finally transitions to an evidence-focused data mixture during learning-rate decay. For supervised fine-tuning (SFT), we regenerate visual question-answering responses using answer stability, answer-masked reconstruction, and correctness-constrained selection to produce more informative and answer-consistent supervision. The 4B model is further optimized with reinforcement learning with verifiable rewards (RLVR). Across five medical multimodal benchmarks, InfiMed2-4B achieves 66.73% mean accuracy after RLVR, surpassing the larger Qwen3.5-9B, while InfiMed2-27B reaches 73.72%, the highest among the evaluated open-weight models.