Related Experiment Video
Updated: May 13, 2026

Systematic Assessment of Mammalian Skull Specimens for Dental and Temporomandibular Joint Pathology
Published on: August 22, 2022
Large language models for the screening step in systematic reviews in dentistry
Rata Rokhshad1, Mobina Bagherianlemraski2, Sarah Sadat Ehsani3
1Department of Pediatric Dentistry, Loma Linda School of Dentistry, Loma Linda, USA.
Chatbots demonstrated limited accuracy and low agreement in screening studies for systematic reviews, particularly in AI-driven tooth segmentation on dental radiographs. Human oversight remains essential for maintaining review integrity.
Area of Science:
- Artificial Intelligence in Radiology
- Medical Informatics
- Systematic Review Methodology
Background:
- Systematic reviews (SRs) are crucial for evidence synthesis in healthcare.
- Automating SR screening steps using artificial intelligence (AI) chatbots is an emerging area.
- Evaluating chatbot performance in specialized tasks like dental radiograph analysis is necessary.
Purpose of the Study:
- To assess the performance of five leading AI chatbots in the screening phase of a systematic review.
- To specifically evaluate their efficacy in identifying studies on tooth segmentation using AI on dental radiographs.
- To compare chatbot screening accuracy and inter-rater agreement against expert human reviewers.
Main Methods:
- A systematic search was conducted across seven major databases for studies on AI-based tooth segmentation in dental radiographs.
- Five chatbots (ChatGPT-4, Claude 2 100k, Claude Instant 100k, LLaMA 3, Gemini) were used to screen retrieved articles.
- Performance was evaluated using accuracy, precision, sensitivity, specificity, F1-score, and Fleiss' Kappa, comparing against expert reviewers.
Main Results:
- Significant variability was found in study inclusion/exclusion rates among chatbots (p < 0.001).
- Claude-instant-100k had the highest inclusion rate (54.88%), while Gemini excluded the most studies (67.90%).
- ChatGPT-4 showed the highest precision (24%) and accuracy (75%), but Fleiss' Kappa indicated systematic disagreement between chatbots.
Conclusions:
- AI chatbots exhibit limited accuracy and low inter-rater agreement in the study screening process for systematic reviews.
- While chatbots can theoretically streamline systematic review tasks, human oversight is critical.
- Maintaining the integrity and reliability of systematic reviews necessitates continued human involvement in the screening phase.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025