Related Experiment Video
Updated: Sep 30, 2026

A Multimodal Imaging Framework to Advance Phenotyping of Living Label-free Breast Cancer Cells
Published on: August 22, 2025
Diagnostic Accuracy of Multimodal Large Language Models (LLMs) for Detecting Tuberous Breast Deformity from
Edoardo Caimi1, Stefano Vaccari2,3, Alessandro Marco Lupacchini1
1Department of Medical Biotechnology and Translational Medicine BIOMETRA, Reconstructive and Aesthetic Plastic Surgery School, University of Milan, Via Festa Del Perdono 7, Milan, Italy.
Background:
Tuberous breast deformity (TBD) is frequently missed outside specialist units, delaying treatment and compounding psychosocial harm. Vision-enabled large language models (LLMs) could facilitate earlier recognition, but their diagnostic accuracy for TBD has not been quantified. The aim of this study was to benchmark four state-of-the-art multimodal LLMs in [1] detecting TBD from standardized photographs and [2] proposing evidence-based surgical strategies.
Methods:
In this single-center, retrospective in-silico study we analyzed 400 photographs (frontal and lateral views) from 200 women aged 16-40 years (100 TBD, 100 controls). GPT-4o, Claude 3 Opus, Gemini 2.5 Pro and DeepSeek V3 were interrogated with a two-tier prompting protocol. Diagnostic accuracy, sensitivity, specificity, PPV, NPV and AUROC were calculated against a surgeon-established reference standard, and 95% CIs were computed by Wilson method. Treatment outputs were independently scored [1-5] by two blinded plastic surgeons; reliability was assessed with the intraclass correlation coefficient (ICC).
Results:
GPT-4o achieved the highest overall accuracy (0.94), sensitivity (0.92) and specificity (0.93); AUROC = 0.88. Claude 3 Opus, Gemini 2.5 Pro and DeepSeek V3 followed with accuracies of 0.90, 0.82 and 0.85, respectively. Mean treatment-quality scores paralleled diagnostic performance: GPT-4o = 4.0 ± 0.4, Claude 3 Opus = 2.8 ± 0.5, Gemini 2.5 Pro = 2.6 ± 0.5, DeepSeek V3 = 2.3 ± 0.5; rater agreement was excellent (ICC = 0.94).
Conclusions:
Among contemporary LLMs, GPT-4o most accurately identifies TBD and delivers the most comprehensive operative plans. With expert oversight, such models could streamline aesthetic breast consultations and resident education.
Level Of Evidence Iii:
Breast Surgery. This journal requires that authors assign a level of evidence to each submission to which Evidence-Based Medicine rankings are applicable. This excludes Review Articles, Book Reviews, and manuscripts that concern Basic Science, Animal Studies, Cadaver Studies, and Experimental Studies. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266.