SentZero improves zero-shot chest x-ray diagnosis with better sentence alignment

SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Understanding chest X-rays using paired images and reports is hard because reports are long and complex. The authors created SentZero, a new method that breaks down reports into meaningful sentences and matches them better with images. This helps the computer recognize medical conditions without needing extra training on specific tasks. SentZero also reduces mistakes caused by repeated or similar sentences. It performs better than earlier methods when tested on different chest X-ray tasks.

What this means in practice

  • For medical imaging developers: Create chest X-ray analysis tools that work without extra training on specific diseases by using improved sentence-image matching.
  • For clinical decision support teams: Enhance diagnostic systems to interpret chest X-rays across various tasks using sentence-focused vision-language models.
  • For automated film reading services: Build more reliable zero-shot X-ray reading services that require less fine-tuning for multi-condition diagnosis.$Commercial implications: Enables sale of multi-disease chest X-ray reading software with zero-shot capabilities for healthcare providers.

Authors

Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho Yang

Abstract

Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.