Vision language method improves acne lesion detection from images

VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Acne severity is hard to measure because traditional ways are subjective and don’t consider how big the pimples actually are. The authors developed a new computer method called VL-AcneSeg that uses both pictures and text clues about face areas to better find and measure acne lesions. Their method works well even on photos taken with smartphones and does not need extra training for new datasets. This could make acne assessment more objective and easier to do outside of clinics.

What this means in practice

  • For dermatology clinics: Provide objective and area-based measures of acne severity using routine clinical photos without manual lesion location input.
  • For mobile health app developers: Integrate acne lesion detection that works reliably on smartphone images to support remote skin condition monitoring.$Commercial implications: Enables consumer apps that offer automated acne analysis and tracking using users’ phone photos.

Authors

Sukju Oh, Soo Ick Cho, Dae Hun Suh, Sukkyu Sun

Abstract

Acne assessment is crucial for clinical decision-making, yet traditional grading and counting are subjective and fail to account for lesion size. While area-based assessment has emerged as a promising alternative, acne segmentation has continued to rely on general-purpose architectures. To address this gap, we propose VL-AcneSeg, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face. Because region-level prompts indicate which facial areas contain lesions, we report a single global prompt, which requires no such information, as our primary setting. On our internal clinical dataset, VL-AcneSeg achieves a Dice score of 0.5082 and an IoU of 0.3407 under this protocol, the highest among all compared methods, including recent vision-language segmentation methods that are themselves given region-level prompts; region-level prompting raises these to 0.5296 and 0.3602. Moreover, lesion area measurements derived from our segmentation correlate with IGA scores at a level comparable to expert annotations (Pearson r = 0.719 versus 0.658). Notably, our framework maintains consistent performance across external validation datasets, performing reliably even on uncontrolled smartphone images without requiring additional training or fine-tuning. By pairing a protocol that requires no lesion-location information with area-based severity estimation, this work provides a foundation for objective acne assessment outside the clinic. Our implementation is publicly available at: https://github.com/sukjuoh/VL-AcneSeg