Related Experiment Video
Updated: Jun 8, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
505
Evidence-Based Potential of Generative Artificial Intelligence Large Language Models on Dental Avulsion: ChatGPT
Taibe Tokgöz Kaplan1, Muhammet Cankar2
1Department of Pedodontics, Faculty of Dentistry, Karabuk University, Karabük, Turkey.
Summary
Gemini demonstrated higher accuracy than ChatGPT in answering dental avulsion questions based on International Society of Dental Traumatology guidelines. Further research is needed to integrate these AI language models into clinical practice.
Area of Science:
- Dental Traumatology
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Dental avulsion, a common dental emergency, requires accurate information for effective management.
- Artificial intelligence (AI) language models offer potential for information dissemination in healthcare.
- Evaluating the accuracy of AI responses in specialized fields like dentistry is crucial.
Purpose of the Study:
- To comparatively evaluate the accuracy and comprehensiveness of answers provided by ChatGPT and Gemini regarding dental avulsion.
- To assess the performance of AI language models against established dental traumatology guidelines.
Main Methods:
- Thirty-three questions on dental avulsion were formulated based on International Society of Dental Traumatology (IADT) guidelines.
- Questions included multiple-choice, binary, and open-ended formats, covering technical and patient queries.
- Responses from ChatGPT and Gemini were scored by four pediatric dentists, with statistical analyses including ICC and Mann-Whitney U tests.
Main Results:
- Gemini achieved a statistically significantly higher mean score than ChatGPT (p=0.001).
- While ChatGPT excelled in open-ended and true/false questions, it showed lower accuracy in multiple-choice questions.
- Gemini's responses showed no significant difference in accuracy across different question types (p=0.088), and overall, Gemini answers were significantly more accurate (p=0.004).
Conclusions:
- Both Gemini and ChatGPT show potential for providing information on dental avulsion, guided by IADT standards.
- Further research, clinical validation, and model improvements are essential for successful integration of these large language models (LLMs) into clinical practice.

