Related Experiment Video
Updated: Oct 10, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
State-of-the-art deep learning for dental implant detection: a standardized benchmark for clinical practice
Jun-Seok Kang1, Yeon-Tae Kim2, Akhilanand Chaurasia3
1Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA, USA.
Purpose:
Although deep-learning approaches show promise for detecting dental implants, the absence of standardized benchmarks has hindered evidence-based model selection. This study systematically compared five contemporary architectures on the same dataset under controlled conditions. A three-seed training-stability analysis was also performed to characterize between-run variability.
Methods:
YOLOv8-Large, YOLOv11-Large, real-time detection transformer (RT-DETR)-Large, lightweight detection transformer (LW-DETR)-Large, and Faster R-CNN ResNet50-FPN were evaluated using 3,996 annotated panoramic radiographs divided into training, validation, and test sets at an 80:10:10 ratio. Each model was trained using three random seeds to assess training stability. Detection accuracy, measured using the F1 score and mean average precision (mAP), end-to-end pipeline speed, measured using frames per second and latency, and graphics processing unit memory use were evaluated. Pairwise statistical comparisons were performed using the independent-samples t-test with Bonferroni correction. To ensure strict comparability among architectures, all mAP values were calculated using a single unified evaluator, pycocotools COCOeval, applied identically to every model.
Results:
YOLOv11-Large achieved the highest mAP@0.5 (0.9873±0.0004) and the fastest end-to-end processing speed (87.2 frames per second), with an F1 score of 0.9853±0.0010. YOLOv8-Large performed similarly, with an mAP@0.5 of 0.9859±0.0005 and an F1 score of 0.9868±0.0002; the 95% confidence intervals for mAP@0.5 overlapped between the two models. RT-DETR-Large had the third-highest accuracy (mAP@0.5 = 0.9806±0.0055). LW-DETR-Large achieved an mAP@0.5 of 0.9699±0.0025 but had substantially slower inference (7.1 frames per second), attributable to quadratic attention scaling at a resolution of 1,536×1,536 pixels. Faster R-CNN ResNet50-FPN had lower accuracy (mAP@0.5 = 0.9214±0.0051) and processing speed (14.2 frames per second). All four contemporary detectors significantly outperformed Faster R-CNN ResNet50-FPN (Bonferroni-corrected P<0.005).
Conclusions:
Within this single-center benchmark based on panoramic radiographs, YOLOv11-Large provided the best balance between processing speed and mAP on workstation-class hardware. LW-DETR-Large achieved competitive accuracy but was constrained by Vision Transformer attention scaling at the clinical image resolution used in this study. These findings may inform model selection for dental implant detection on panoramic radiographs. External multicenter and cross-modality validation is required before broader claims regarding clinical deployment can be made.
