Scientific image quality assessed by multimodal retrieval system
Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation
Computer Vision and Pattern RecognitionComputation and Language
Summary
Judging the quality of scientific images can be tricky because it requires understanding both the pictures and their scientific context. The authors created a system that helps a language model by giving it relevant visual and text examples to compare with the images being evaluated. This approach helps the system better think like human experts when assessing image quality. Their method won first place in a scientific image quality assessment challenge.
What this means in practice
- •For scientific imaging teams: Automatically assess and score the quality of complex scientific images in research workflows using multimodal reference retrieval techniques.
- •For medical imaging analysts: Assist evaluation of medical imaging quality by referencing multimodal examples to improve diagnostic image credibility.
Authors
Yinuo Zhang, Bingshuo Liu, Zhiying Tu, Dianhui Chu, Qingbin Liu, Xi Chen, Jiang Bian, Xiaoyan Yu, Dianbo Sui
Abstract
This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.