Related Experiment Video
Updated: Sep 8, 2025

Effects of Mechanical Methods Used in Peri-implantitis Treatment on Implant Surface Decontamination and Roughness
Published on: March 14, 2025
Assessing the diagnostic and treatment accuracy of Large Language Models (LLMs) in Peri-implant diseases: A clinical
Igor Amador Barbosa1, Mauro Sergio Almeida Alves1, Paloma Rayse Zagalo de Almeida1
1Dental Clinic Post-Graduate Program, University Center of State of Pará, Belém, Pará, Brazil.
Objective:
This study evaluated the coherence, consistency, and diagnostic accuracy of eight AI-based chatbots in clinical scenarios related to dental implants.
Methods:
A double-blind, clinical experimental study was carried out between February and March 2025, to evaluate eight AI-based chatbots using six fictional cases simulating peri‑implant mucositis and peri‑implantitis. Each chatbot answered five standardized clinical questions across three independent runs per case, generating 720 binary outputs. Blinded investigators scored each response against a gold standard. Statistical analyses included chi-square and Fisher's exact and Cohen's Kappa tests were used to assess intra-model consistency, stability and reliability for each AI chatbot.
Results:
GPT-4o demonstrated the highest diagnostic accuracy (88.8 %), followed by Gemini (77.7 %), OpenAI o3-mini (72.2 %), OpenAI o3-mini-high (71.1 %), Claude (66.6 %), OpenAI o1 (60 %), DeepSeek (55.5 %), and Copilot (49.9 %). GPT-4o also showed the highest intra-model stability (κ = 0.82) and consistency, while Copilot and DeepSeek showed the lowest reliability. Significant differences were observed only in the reference citation criterion (p < 0.001), with Gemini being the only AI chatbot to achieve 100 % compliance, but GPT-4o consistently outperformed the other AI chatbots across all evaluation domains.
Conclusion:
GPT-4o demonstrated superior diagnostic accuracy and response consistency, reinforcing the influence of AI chatbot architecture and training on clinical reasoning performance. In contrast, Copilot showed lower reliability and higher variability, emphasizing the need for cautious, evidence-based adoption of AI tools in the diagnosis of peri‑implant diseases.
Clinical Relevance:
Understanding AI performance in peri‑implant diagnosis to support evidence-based decision-making using AI and its responsible clinical use.

