Related Experiment Video
Updated: Jun 13, 2026

Guided Endodontics: Three-Dimensional Planning and Template-Aided Preparation of Endodontic Access Cavities
Published on: May 24, 2022
Evaluating the Ability of Multimodal Artificial Intelligence to Identify Endodontic Instruments: A Comparative Study
1Department of Endodontics, Faculty of Dentistry, Pamukkale University, 20160 Denizli, Türkiye.
None:
Background/Objectives: Multimodal large language models (LLMs) are increasingly integrated into dental diagnostics. This study evaluated the ability of ChatGPT-4o and Gemini 3 Flash to visually identify endodontic instruments and assess their explanatory plausibility regarding instrument morphology. Methods: Standardized images of five endodontic file systems (Reciproc R25, Reciproc Blue, WaveOne Gold, MM One Shape, and XP-endo Finisher) were submitted to both models via their free tiers. Each image was evaluated 50 times per model (total n = 500) to assess both classification accuracy and response consistency. Visual recognition performance was measured using recall, precision, and F1-score, while the plausibility of morphological explanations was evaluated using a structured 3-point scale. Results: Gemini 3 Flash demonstrated significantly higher recognition performance compared to ChatGPT-4o (p < 0.001). The overall acceptable response rate was higher for Gemini 3 Flash (94.4%, [95% CI: 91.5-97.3%]) than for ChatGPT-4o (67.2%, [95% CI: 61.4-73.0%]; p < 0.001). Notably, Gemini 3 Flash showed strong performance in identifying complex instrument designs, whereas ChatGPT-4o exhibited marked limitations in recognizing certain non-standard geometries. Reliability analysis indicated higher consistency for Gemini 3 Flash (κ = 0.86, [95% CI: 0.81-0.91]) compared to ChatGPT-4o (κ = 0.51, [95% CI: 0.44-0.58]). Conclusions: Gemini 3 Flash outperformed ChatGPT-4o in both classification accuracy and consistency in this controlled visual identification task. While these findings highlight the potential of multimodal LLMs in endodontic workflows, their current performance variability limits direct, autonomous clinical application. Further validation under clinically realistic conditions is required before such systems can be considered reliable adjunctive tools.
