Medical system models patient changes from images and reports
Clinical Trajectory Alignment for Medical Vision-Language Pre-training
Computer Vision and Pattern Recognition
Summary
Medical imaging and report data usually get matched one visit at a time, which misses how a patient's condition changes over many visits. The authors propose MedCTA, a method that looks at changes in specific health issues and the overall patient history across time. It uses artificial intelligence to read many reports automatically and learns how conditions evolve without needing people to label the time-based changes. This approach helps computers better track disease progress and improves tasks like classifying images and matching images with reports.
What this means in practice
- •For hospital data teams: Enhance patient monitoring systems by integrating models that track imaging and report changes over time to improve detection of clinical progression.
- •For medical imaging software developers: Build AI tools that better match longitudinal patient images and reports, enabling more precise diagnostics and treatment planning.
Authors
Huimin Yan, Xian Yang, Zhi Wang, Liang Bai
Abstract
Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.