Flow matching improves speed and detail in SAR to optical image conversion
ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning
Computer Vision and Pattern Recognition
Summary
Turning radar images (SAR) into regular photos that people can understand is hard because details often get blurry or lost in the process. The authors created a new method called ContraFM-S2O that uses a mathematical technique called flow matching to convert images faster and with clearer details. Instead of taking many slow steps, their approach predicts the full transformation in one go. They also use a learning method called contrastive learning to make the converted images more accurate and sharp. Tests show their method works better and faster than older methods on common datasets.
What this means in practice
- •For remote sensing teams: Generate clear optical images from SAR data in one step for faster environmental monitoring and analysis.
- •For satellite operators: Produce high-fidelity optical images from SAR satellite data quickly to support real-time decision making.
Authors
Mingqian Yu, Wei-kuan Chiang, Qiurui Wang, Peilin Zhao
Abstract
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image translation. Unlike conventional diffusion models, ContraFM-S2O learns to predict the velocity field in training and solves ODE instead of SDE during inference to improve the sampling efficiency. In addition, ContraFM-S2O replaces instantaneous velocity with average velocity along the interpolation path to realize one-step SAR-to-optical image translation and uses contrastive learning to improve the quality of the generated optical images. Experiments show our model achieves state-of-the-art on SAR2Opt and QXS datasets, outperforming baselines, and reduces inference latency via one-step generation.