Related Experiment Video
Updated: Apr 28, 2026

05:49
Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
1.7K
Prognostic Prediction of Avulsed Permanent Teeth Using Conversational AI Models Versus Expert Dentists: Influence of
Mehmet Buldur1, Gizem Ayan1, Tugba Misilli1
1Department of Restorative Dentistry, Faculty of Dentistry, Çanakkale Onsekiz Mart University, Çanakkale, Türkiye.
Summary
Conversational AI models show meaningful agreement with dentists for predicting tooth avulsion prognosis. Performance varies by model and prompt format, suggesting AI can be a supportive tool for clinicians.
Area of Science:
- Dental Prognosis
- Artificial Intelligence in Dentistry
- Clinical Decision Support Systems
Background:
- Prognosis prediction for tooth avulsion is complex, influenced by multiple clinical factors.
- The reliability and consistency of conversational large language models (LLMs) in this domain are not well-established.
- Understanding LLM sensitivity to prompt variations and output stability is crucial for potential clinical integration.
Purpose of the Study:
- To compare the prognostic accuracy of four conversational LLMs against expert dentist assessments for tooth avulsion.
- To evaluate the impact of different prompt formats on LLM performance in this context.
- To assess the short-term stability of LLM outputs for avulsion prognosis.
Main Methods:
- A simulation-based study using 120 standardized tooth avulsion scenarios.
- Four LLMs (ChatGPT-4o, Gemini, Claude, DeepSeek) evaluated scenarios in four prompt formats, twice over 48 hours.
- Expert dentist consensus served as the reference standard for agreement, variability, and stability analyses.
Main Results:
- LLMs demonstrated meaningful agreement with expert prognosis predictions for both numeric scores and ordinal categories.
- Gemini and ChatGPT-4o exhibited balanced performance; Claude maintained risk ordering but deviated in severity; DeepSeek had lower categorical concordance.
- Prompt format significantly affected LLM outputs, with clinical-practical formats yielding better alignment than narrative ones. Short-term stability was generally good, with minor shifts in some combinations.
Conclusions:
- Conversational LLMs can provide avulsion prognosis estimates that align with expert dentists under controlled conditions.
- Model performance is dependent on the specific AI and the structure of the input information.
- LLMs show potential as secondary decision aids for clinicians in prognosis estimation, but real-world data is needed to confirm utility and stability.

