Related Experiment Video
Updated: Feb 28, 2026

05:49
Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
1.6K
Artificial Intelligence Versus Human Dental Expertise in Diagnosing Periapical Pathosis on Periapical Radiographs: A
Fatma E A Hassanein1, Radwa R Hussein2, Mohamed Riad Elgarhy3
1Oral Medicine, Periodontology, and Oral Diagnosis, Faculty of Dentistry, King Salman International University, El Tur 46612, Egypt.
Bioengineering (Basel, Switzerland)
|February 27, 2026
Summary
Artificial intelligence (AI) shows high sensitivity but low specificity in detecting periapical pathosis on radiographs, tending to over-diagnose. Dental professionals must validate AI findings for accurate clinical diagnosis.
Area of Science:
- Dental Radiology
- Artificial Intelligence in Medicine
- Diagnostic Accuracy
Background:
- Periapical pathosis diagnosis is crucial for endodontic success but challenged by 2D imaging limitations and interpretation subjectivity.
- Artificial intelligence (AI) presents a potential solution, yet its diagnostic granularity compared to human clinicians requires thorough investigation.
- Evaluating AI's role in interpreting periapical radiographs is essential for advancing diagnostic capabilities.
Purpose of the Study:
- To assess the diagnostic accuracy of ChatGPT-5 in identifying periapical radiographic abnormalities.
- To compare the performance of ChatGPT-5 against a consensus of three board-certified oral radiologists.
- To investigate the agreement and predictors of correct AI classification in periapical radiograph interpretation.
Main Methods:
- A retrospective diagnostic accuracy study involving 270 periapical radiographs.
- Independent interpretation of radiographs by ChatGPT-5 using a standardized prompt and by a three-expert oral radiologist consensus.
- Comparison of diagnostic accuracy, agreement (Cohen's kappa), and classification predictors between AI and expert consensus.
Main Results:
- ChatGPT-5 achieved high sensitivity (87.5%) but low specificity (12.5%), yielding 50.0% overall diagnostic accuracy.
- The AI model exhibited a tendency to over-identify pathology, classifying 87.5% of radiographs as abnormal versus 50.0% by experts.
- Agreement was near-perfect for anatomical localization (κ = 0.857) but poor for abnormality detection (κ = 0.000); significant disagreement noted for lesion border characterization (κ = 0.127).
Conclusions:
- ChatGPT-5 demonstrates high sensitivity in visual interpretation of periapical radiographs but limited specificity and inconsistent morphological feature characterization hinder its clinical diagnostic reliability.
- The AI system systematically over-diagnoses and describes lesions differently than dental experts, categorizing them as more structurally defined.
- AI holds potential as a sensitive initial screening tool, but findings necessitate validation by dental professionals to mitigate false positives and ensure accurate characterization.

