Related Experiment Videos
Guideline-Based Evaluation of Large Language Models in Psoriasis Treatment
Kuanyu Xia1,2,3, Lang Min4, Dan Jian1,2,3
1Department of Dermatology, Xiangya Hospital, Central South University, Changsha, China.
Background:
The therapeutic landscape for moderate-to-severe psoriasis has expanded rapidly with biologics and small molecules, increasing the complexity of clinical decision-making. Although large language models (LLMs) may support medical information synthesis, their reliability in high-stakes dermatologic management remains uncertain.
Objectives:
To compare the guideline adherence, safety, and clinical decision quality of multiple LLMs with dermatology experts in systemic psoriasis management.
Methods:
A 40-scenario question bank was developed from the Living EuroGuiDerm Guideline for the Systemic Treatment of Psoriasis Vulgaris and stratified by risk level. Five LLMs (ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3.0 Pro, Grok 4.1 Thinking, and DeepSeek V3.2) were benchmarked against five board-certified dermatologists. Responses were scored for completeness and safety.
Results:
Normalized completeness scores differed significantly among study groups (χ2 = 20.55, df = 5, P < 0.001, ε2 = 0.066). Gemini 3.0 Pro, DeepSeek V3.2, and Grok 4.1 Thinking scored significantly higher than the expert panel. This advantage was driven primarily by must-know items, for which hit rates differed significantly among groups (P = 0.003; 68% for the human benchmark vs 79-90% for LLMs). In low-risk scenarios, AI models outperformed experts (χ2 = 17.54, P = 0.004), whereas high-risk performance converged (χ2 = 6.14, P = 0.293). However, safety rankings reversed on critical error analysis (Q = 17.10, df = 5, P = 0.004): experts made 1 critical error, while LLMs made 5-9 each. Experts made no critical errors in high-risk scenarios, whereas all LLMs made at least one.
Conclusions:
Current LLMs can match or exceed dermatologists in completeness of guideline-based systemic psoriasis recommendations, particularly in low-risk contexts. However, they remain more prone to critical safety errors in complex, high-stakes scenarios.