Related Experiment Video
Updated: Sep 5, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Evaluation of LLMs for Automated Dental-to-Medical Consult Letter Generation: Comparative Study
Dareen Eom1, Kyeol Koh2, Howon Chung3
1Interdisciplinary Program of Medical Informatics, College of Medicine, Seoul National University, Seoul, Republic of Korea.
Introduction And Aims:
For dental patients with systemic diseases, multidisciplinary care via referral letters is essential for clinical safety. This study evaluated the clinical feasibility and quality of dental-to-medical referral letters generated by Large Language Models (LLMs).
Methods:
Using 153 paired datasets, three LLMs - Gemini-2.5-Pro, GPT-5 and MedGemma - generated referral letters via a standardised prompting framework. Two dental clinicians performed a blinded assessment across nine sub-criteria covering content, style and correctness. GPT-5 also performed an automated evaluation. Statistical significance was analysed using Friedman and Wilcoxon signed-rank tests with Holm correction.
Results:
In the human clinical evaluation, Gemini-2.5-Pro achieved the highest overall performance (TOTAL mean SCORE: 2.9487 with a bootstrap 95% confidence interval [CI] of 2.9345-2.9606), significantly outperforming GPT-5 (2.8818 with 95% CI of 2.8613-2.9008) and MedGemma (2.8613 with 95% CI of 2.8338-2.8865) (P < .05). While GPT-5 showed high clarity in systemic problem descriptions, MedGemma excelled in structural adherence. The overall performance difference between GPT-5 and MedGemma was not statistically significant (P = .3519). Inter-rater reliability was strong in the prompt search dataset (QWK: 0.7175) but lower in the independent test dataset (QWK: 0.1505), reflecting the subjective nature of clinical document assessment; therefore, final human scores were reported as average ratings from two blinded clinicians with bootstrap confidence intervals. Notably, GPT-5's self-evaluation results deviated from human judgment, exhibiting potential model-based bias.
Conclusion:
This preliminary study suggests the potential of Gemini-2.5-Pro for automated consult letter generation, though the low test-set inter-rater reliability (QWK: 0.1505) highlights the subjective complexity of clinical text. To integrate these assistive prototypes into routine clinical application, more systematic studies are required.
Clinical Relevance:
Integrating LLMs offers a promising approach to assisting with dental-to-medical referrals, presenting a potential future benefit of reducing administrative workloads and enabling healthcare providers to focus on safe, empathetic patient care.
