Related Experiment Video
Updated: May 13, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.3K
Evaluating Large Language Models for Burning Mouth Syndrome Diagnosis
Takayuki Suga1, Osamu Uehara2, Yoshihiro Abiko3
1Department of Psychosomatic Dentistry, Graduate School of Medical and Dental Sciences, Institute of Science Tokyo, Tokyo, Japan.
Journal of Pain Research
|March 24, 2025
Summary
Large language models show promise for diagnosing burning mouth syndrome, with ChatGPT and Claude achieving 99% accuracy. Clinician oversight remains crucial due to model variations and potential errors in diagnosis.
Area of Science:
- Artificial Intelligence in Dentistry
- Medical Diagnostics
- Oral Medicine
Background:
- Burning mouth syndrome presents diagnostic challenges due to its subjective nature and lack of clear etiology.
- Large language models (LLMs) are emerging as potential tools for medical diagnosis across various specialties.
Purpose of the Study:
- To evaluate the diagnostic accuracy of leading LLMs in identifying burning mouth syndrome.
- To explore potential limitations and variations in LLM diagnostic reasoning for this condition.
Main Methods:
- 100 synthesized clinical vignettes of burning mouth syndrome cases were assessed by three LLMs: ChatGPT-4o, Gemini Advanced 1.5 Pro, and Claude 3.5 Sonnet.
- LLMs were prompted for primary diagnosis, differential diagnoses, and reasoning, with accuracy compared against expert evaluations.
Main Results:
- ChatGPT and Claude demonstrated 99% accuracy, significantly outperforming Gemini Advanced (89% accuracy, p < 0.001).
- Observed misdiagnoses included Persistent Idiopathic Facial Pain and other conditions, with notable differences in reasoning patterns and data requests among models.
Conclusions:
- LLMs show significant potential as supplementary diagnostic aids for burning mouth syndrome, particularly in resource-limited settings.
- Clinician oversight and expert verification are essential due to observed variations in model performance and reasoning, ensuring accurate patient management.
Keywords:
artificial intelligenceburning mouth syndromedentistrydiagnostic accuracylarge language models
