Multimodal AI models perform near dermatologists on patient skin images

Benchmarking Off-the-Shelf Multimodal AI Models Against Dermatologists on Patient-Captured Skin Images

Computer Vision and Pattern Recognition

Summary

Diagnosing skin conditions from photos can be hard, even for doctors. This paper tests three affordable AI models to see how well they can identify skin problems from patient-taken pictures. The researchers found that these AI models perform close to the level of certified dermatologists but have issues like unreliable confidence ratings. Also, adding patient information did not always help, and paying more for an AI model didn’t guarantee better results.

What this means in practice

  • For dermatology clinics: Use cost-effective AI tools to assist in diagnosing skin conditions from patient photos without expecting reliable confidence scores from the AI.
  • For healthcare app developers: Integrate affordable multimodal AI models into mobile apps to support remote skin condition assessments from user-submitted photos with caution about patient data effects.$Commercial implications: Enables selling AI-powered teledermatology features that automate preliminary skin diagnosis from mobile images.

Authors

Rian Dolphin, Laura Knowles

Abstract

Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attention. However, recent generations of frontier AI models have adopted multimodal capabilities as a first class citizen, with vision capabilities being central to that. In this paper, we evaluate three recently released models on the task of diagnosing dermatological conditions from patient-submitted images. The models chosen are at the low to mid tier in terms of pricing and thus represent a floor on current AI capabilities, not a ceiling. We evaluate AI performance relative to a panel of three certified dermatologists, who grade each image, and we present four interesting findings. Firstly, depending on the metric, the tested AI models are either on par or slightly trail humans in terms of inter-clinician agreement. Secondly, we find that asking AI models for a confidence rating produces poorly calibrated answers, meaning use of confidence thresholds should not be relied upon in a clinical setting. Thirdly, the effect of providing additional patient metadata is strongly model-specific, with one of the three models degrading on every metric considered. Finally, model cost is not predictive of performance. The best-performing model we tested costs on average $0.0045 per case.