Related Experiment Video
Updated: Sep 16, 2026

Granulocyte-dependent Autoantibody-induced Skin Blistering
Published on: October 12, 2012
Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering
Asmaa Abou-Bakr1, Salma M Saad2, Nevine H Kheir El Din3
1Oral Medicine and Periodontology, Faculty of Dentistry, Galala University, Suez 43511, Egypt.
Abstract:
Background/Objectives: To compare the diagnostic performance of Claude Opus 4.7 and Gemini Pro 3 for the differential diagnosis of oral autoimmune blistering diseases (AIBDs) and evaluate their diagnostic reasoning, confidence, and calibration. Materials and Methods: This retrospective multicenter paired diagnostic accuracy study included 200 clinicopathologically confirmed AIBD cases (50 each of pemphigus vulgaris, mucous membrane pemphigoid, bullous pemphigoid, and linear IgA bullous dermatosis). Each case was independently assessed by both models using identical standardized clinical information and clinical photographic inputs. The task required forced-choice classification among the four predefined diseases. Histopathological and direct immunofluorescence findings were used exclusively to establish the clinicopathological reference diagnosis and were not provided to the AI models. The reference diagnosis was established by clinicopathological correlation. The primary outcome was diagnostic accuracy. Secondary outcomes included disease-specific diagnostic performance, Cohen's κ, ROC analysis, calibration, confidence, reasoning quality, management recommendations, and error patterns. Pre-consensus inter-rater reliability of the two human assessors was also evaluated using Cohen's κ for binary outcomes and weighted Cohen's κ for the ordinal reasoning-quality score. Results: Claude achieved significantly higher diagnostic accuracy than Gemini (92.0% vs. 86.0%, p = 0.012), stronger agreement with the reference standard (κ = 0.893 vs. 0.813), and superior discrimination (macro-AUC 0.998 vs. 0.965). Claude demonstrated higher key diagnostic-feature identification (92.0% vs. 86.0%; p = 0.012) and higher clinical-reasoning scores (61.0% vs. 40.0% of responses rated good; Wilcoxon p < 0.001; r = 0.47), whereas management recommendations did not differ significantly (100.0% vs. 98.0%; p = 0.125). Calibration results were metric-dependent: Claude had a lower one-vs-rest Brier score (0.0468 vs. 0.0558), whereas Gemini had a lower expected calibration error (0.083 vs. 0.251). For both models, the predominant error was misclassification of linear IgA bullous dermatosis as mucous membrane pemphigoid. Conclusions: Both multimodal LLMs showed high performance in this controlled four-class benchmark, with Claude Opus 4.7 outperforming Gemini Pro 3 in overall accuracy and reasoning quality. However, these findings do not establish autonomous diagnostic capability, clinical effectiveness, or safety. The LABD-MMP misclassification and metric-dependent calibration highlight important limitations. The models should therefore be regarded as investigational adjunctive decision-support tools requiring clinician oversight and diagnostic verification. Prospective external and human-in-the-loop validation is required before clinical implementation.