Summary
Predicting how physical systems behave often requires combining different types of data and knowledge. Existing AI methods treat physical measurements and equations as simple numbers or text, missing important physical rules that guide how these parts interact in space and time. This work introduces a new way to represent physical data called a physical-field modality, which respects these physical rules. The researchers built PhyMo, a system that first learns to understand physical fields, then aligns this knowledge with visual data, and finally uses the combined information to make better predictions. Tests on multiple physics problems show that PhyMo outperforms previous methods in understanding and predicting physical phenomena.
What this means in practice
- •For engineering simulation teams: Improve simulation accuracy by integrating sensor measurements with physical field models for better prediction of physical systems behavior.
- •For climate modelers: Enhance climate predictions by combining satellite imagery with physics-based field representations in a unified AI framework.
Authors
Henan Sun, Haitao Hu, Jin Liu, Jianfeng Zhang, Lujia Pan, Nuo Chen, Jia Li
Abstract
Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.