Related Experiment Video
Updated: Jun 15, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
523
Evaluating the Accuracy, Reliability, Consistency, and Readability of Different Large Language Models in Restorative
Zeyneb Merve Ozdemir1, Emre Yapici1
1Department of Restorative Dentistry, Faculty of Dentistry, Kahramanmaras Sutcu Imam University, Kahramanmaras, Turkey.
Summary
Artificial intelligence (AI) chatbots show varied reliability in answering Restorative Dentistry questions. ChatGPT-4o and Chatsonic offer promising accuracy for patient education, though readability needs improvement.
Area of Science:
- Artificial Intelligence in Dentistry
- Digital Health Technologies
- Restorative Dentistry Applications
Background:
- The integration of artificial intelligence (AI) is rapidly transforming various fields, including dentistry.
- Evaluating the performance of AI chatbots in specialized medical domains is crucial for their safe and effective implementation.
- Restorative dentistry, with its complex information needs, presents a unique area for AI assessment.
Purpose of the Study:
- To assess the reliability, consistency, and readability of responses from multiple AI chatbots to Restorative Dentistry queries.
- To compare the performance of different AI models, including ChatGPT-3.5, ChatGPT-4, ChatGPT-4o, Chatsonic, Copilot, and Gemini Advanced.
- To determine the suitability of AI-generated content for both dental professionals and patients.
Main Methods:
- Forty-five knowledge-based and 20 specific questions (patient-related and dentistry-specific) were posed to six AI chatbots.
- Reliability was evaluated using the DISCERN questionnaire.
- Readability was assessed via Flesch Reading Ease and Flesch-Kincaid Grade Level scores.
- Accuracy and consistency were determined by analyzing responses to knowledge-based questions.
Main Results:
- ChatGPT-4, ChatGPT-4o, Chatsonic, and Copilot demonstrated "good" reliability, while ChatGPT-3.5 and Gemini Advanced had "fair" reliability.
- ChatGPT-4o achieved the highest accuracy (93.3%) for knowledge-based questions.
- Readability scores did not significantly differ among chatbots, but generally exceeded recommended levels for patient education.
- ChatGPT-4 showed the highest consistency in responses.
Conclusions:
- AI chatbot performance in Restorative Dentistry varies significantly in accuracy, reliability, and consistency.
- ChatGPT-4o and Chatsonic show potential for academic and patient education in Restorative Dentistry.
- Current AI response readability may require adjustment for optimal patient comprehension in dental education.
Related Concept Videos
Improving Translational Accuracy
2.5K
2.5K
Data Validation
5.0K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.0K

