Related Experiment Video
Updated: Jun 12, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Large Language Models and Retrieval-Augmented Platforms for Dental Implant Complications: A Blinded, Expert-Rated
Yaniv Mayer1, Luigi Canullo2, Marton Kivovics3
1The Ruth and Bruce Rappaport Faculty of Medicine, Technion - Israel Institute of Technology, Haifa, Israel; Department of Periodontology, Rambam Health Care Campus, Haifa, Israel.
A medical-domain retrieval-augmented (RAG) platform, OpenEvidence, demonstrated superior performance in diagnostic accuracy and safety compared to general-purpose large language models (LLMs) for dental implant complications. Specialist verification remains crucial for all AI tools in clinical practice.
Area of Science:
- Artificial Intelligence in Dentistry
- Medical Informatics
- Clinical Decision Support Systems
Background:
- Conversational artificial intelligence (AI), including large language models (LLMs), is increasingly explored for clinical applications.
- Evaluating the diagnostic accuracy, patient safety, and reliability of AI platforms in complex medical scenarios like dental implant complications is critical.
Purpose of the Study:
- To compare the diagnostic accuracy, patient safety, completeness of management, and hallucination rates of eleven AI platforms using synthetic dental implant complication vignettes.
- To assess the performance differences between general-purpose LLMs, a general-purpose retrieval-augmented (RAG) platform, and a specialized medical-domain RAG platform.
Main Methods:
- Thirty synthetic vignettes covering surgical, biological, and mechanical/prosthetic dental implant complications were used.
- Four blinded dental specialists rated AI responses on a 5-point Likert scale for accuracy, safety, completeness, and hallucinations.
- Statistical analyses included Friedman tests for ordinal scores and Cochran's Q for binary outcomes (safety and hallucination-free rates).
Main Results:
- The medical-domain RAG platform (OpenEvidence) achieved 100% safety and 100% hallucination-free responses across all vignettes.
- General-purpose platforms showed variable performance, with safe rates from 56.7% to 86.7% and hallucination-free rates from 36.7% to 93.3%.
- OpenEvidence demonstrated the highest composite score (4.90 ± 0.18), with significant performance advantages, particularly in mechanical/prosthetic cases.
Conclusions:
- The medical-domain RAG platform (OpenEvidence) outperformed general-purpose LLMs and a general-purpose RAG platform in expert-rated diagnostic accuracy and safety for dental implant complications.
- Retrieval-augmented platforms grounded in peer-reviewed literature showed higher expert ratings than general-purpose LLMs.
- All AI platforms should be used as adjuncts to specialist judgment, requiring prospective validation before clinical deployment.

