Related Experiment Video
Updated: Sep 23, 2026

A Robust Discovery Platform for the Identification of Novel Mediators of Melanoma Metastasis
Published on: March 8, 2022
Evaluating LLMs in non-metastatic melanoma care: a comparative analysis
Xizhi Wu1, Wangyan Zhong1, Wanli Ye1
1Department of Radiation Oncology, Shaoxing People's Hospital, The First Hospital of Shaoxing University, Shaoxing, Zhejiang, China.
Background:
Malignant melanoma is an aggressive skin cancer with rising incidence. Accurate early staging and standardized treatment are crucial for prognosis. This study evaluates seven large language models (LLMs)-including GPT-5.2, Gemini-3.1, and medically enhanced models-in assisting non-metastatic melanoma management amid clinical complexities.
Methods:
Employing a prospective, simulated expert-blinded design, 59 virtual cases across TNM stages, ages, and comorbidities were assessed. Multiple senior oncologists independently evaluated model outputs using a 6-point Likert scale for staging accuracy, treatment rationality, and protocol standardization.
Results:
GPT-5.2 (5.56 ± 1.12) and Gemini-3.1 (5.25 ± 1.3) achieved the highest staging accuracy, while AntAngelMed performed worst (2.93 ± 1.67). Performance declined significantly in complex Stage III cases. GPT-5.2 and Gemini-3.1 also led in treatment rationality, showing stability, whereas model performances converged in early stages but diverged in advanced ones. Gemini-3.1 excelled in protocol standardization (5.17 ± 0.57), though some models posed risks like insufficient surgical margin recommendations.
Conclusion:
Leading LLMs demonstrate potential for high-quality melanoma management but exhibit inconsistent performance influenced by architecture and case complexity, with reduced reliability in advanced stages. Future tools require risk-stratified guidelines and real-world validation to improve patient outcomes.
