Related Experiment Video
Updated: Jun 27, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Diagnostic Performance and Error Patterns of a Large Language Model and Neural Network in Periodontitis
Agata Ossowska1, Aida Kusiak1, Albert Camlet1
1Department of Periodontology and Oral Mucosa Diseases, Medical University of Gdansk, 80-200 Gdansk, Poland.
Journal of Clinical Medicine
|June 26, 2026
Summary
A neural network outperformed a large language model in diagnosing periodontitis staging and grading. The neural network achieved higher accuracy, while the large language model tended to underestimate disease severity.
Area of Science:
- Periodontology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Periodontitis is a common chronic disease necessitating precise diagnosis for treatment.
- Artificial intelligence (AI) offers potential for clinical decision support in periodontitis management.
Purpose of the Study:
- To compare the diagnostic performance of a large language model (LLM) and a neural network (NN) for periodontitis classification.
- To analyze classification error patterns between the LLM and NN using the current staging and grading system.
Main Methods:
- Retrospective analysis of 110 patients with periodontal disease, including clinical and demographic data.
- Evaluation of LLM and NN diagnostic performance using accuracy, confusion matrices, and Cohen's kappa coefficient.
- Comparison of error patterns, including disease severity underestimation/overestimation by each AI model.
Main Results:
- The neural network achieved higher accuracy (85% for stage, 79% for grade) compared to the LLM (62% for stage, 63% for grade).
- The LLM frequently underestimated disease severity, while the NN tended to overestimate progression.
- Statistical analysis confirmed significant performance differences between the LLM and NN (p < 0.0001).
Conclusions:
- A task-specific neural network demonstrated superior diagnostic performance over the evaluated large language model for periodontitis classification.
- Differences in AI performance are linked to distinct training paradigms and intended applications.
- Further research is essential to refine and validate AI tools for clinical periodontitis diagnosis.
