Related Experiment Videos
Large language models in Hidradenitis suppurativa: strong empathy, weak exposome awareness
Ahmet Uğur Atılan1, Niyazi Çetin1, Ahmet Kağan Özdemir2
1School of Medicine, Department of Dermatology, Pamukkale University, Denizli, Türkiye.
Background:
Hidradenitis suppurativa (HS) imposes a substantial biopsychosocial burden. As patients increasingly consult large language models (LLMs) for health information, their capacity to provide clinically appropriate, empathic, readable and exposome-aware counseling requires evaluation.
Objectives:
To benchmark LLM-generated responses to simulated HS consultations for clinical performance, exposome awareness, quality-of-life (QoL) coverage and readability.
Methods:
In this blinded cross-sectional study, five LLMs (Claude-4.5 Sonnet, Gemini-3.0 Pro, ChatGPT-5.2, DeepSeek-V3.2 and Llama-4.0 Maverick 17B) generated zero-shot, single-turn responses to 47 clinically and socially complex HS cases. Three independent dermatologists rated responses across six domains: clinical accuracy, understandability, shared decision-making, actionability, exposome awareness and clinical empathy. The primary outcome was the mean six-domain Likert score, rescaled to a 20-100 normalized score. Secondary outcomes included domain scores, QoL coverage, readability and response length.
Results:
Overall, 235 responses yielded 705 dermatologist evaluations (ICC[3,k] = 0.814; 95% CI, 0.797-0.831). Claude achieved the highest normalized score (87.00), followed by Gemini (81.25), ChatGPT (79.22), DeepSeek (79.20) and Llama (70.61). Clinical empathy (91.09) and understandability (90.16) scored highest, and exposome awareness lowest (52.94; p < 0.0001). Sleep, environmental factors and work/academic performance were rarely addressed. Readability was more demanding than recommended for patient education materials.
Conclusions:
LLMs provide generally empathic and clinically plausible HS counseling but show gaps in exposome awareness, QoL integration and health-literacy-sensitive communication. Within the constraints of single, zero-shot responses to simulated scenarios, these findings support a cautious, clinician-supervised adjunctive role and the need for dermatology-specific optimization before unsupervised patient-facing use.