MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI
2026-08-31 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial Intelligence
AI summaryⓘ
The authors developed MR-JEPA, a new deep learning model that looks at heart MRI videos and images in 3D over time, instead of just single 2D slices like before. They trained it without needing labeled data on many types of MRI sequences from over 10,000 patients. When tested on different heart measurement and disease detection tasks, MR-JEPA performed better than earlier models, especially on measuring heart function and muscle strain. This shows the model can use different MRI views together to help doctors analyze heart health more accurately.
Cardiac Magnetic Resonance Imaging (CMR)self-supervised learningvideo foundation modelcine MRIlate gadolinium enhancement (LGE)myocardial strainejection fractionspatiotemporal datatokenizationattention architecture
Authors
Athira J. Jacob, Puneet Sharma, Dorin Comaniciu, Daniel Rueckert
Abstract
Cardiac magnetic resonance imaging (CMR) produces rich sequential data such as temporal cine videos and spatial LGE/mapping stacks, yet most deep learning approaches process individual 2D slices, discarding this context. We present MR-JEPA, a self-supervised video foundation model for CMR that extends LeJEPA to 3D spatiotemporal inputs through tubelet tokenization, spatiotemporal masking augmentation, and initialization from a 2D CMR foundation model. Unlike prior CMR video models limited to cine data, MR-JEPA is pretrained on multi-sequence data (cine, LGE, mapping) from 10,505 patients across two centers without annotations. We evaluate the frozen encoder on six downstream tasks using a unified multi-view gated attention architecture: LV ejection fraction, RV ejection fraction, three myocardial strains (GLS, GCS, GRS), and four-class disease detection. MR-JEPA outperforms other compared methods on all five regression tasks, including both a domain-specific CMR model pretrained on more data with text supervision and a natural-video foundation model, achieving an LV EF MAE of 4.79% (r =0.764) and a GLS MAE of 1.87 (r=0.805), with 21-27% MAE reductions over baselines on strain tasks. For disease detection, MR-JEPA achieved a macro AUG of 0.868, remaining competitive with the domain-specific baseline despite using a fully self-supervised pretraining objective. These results demonstrate the potential of a unified video encoder for robust, multi-view utilization of diverse CMR sequences in clinical cardiac quantification and diagnosis.