Related Experiment Video
Updated: Feb 28, 2026

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
CBCT-Based Orthodontic Classification Using Commercial AI: Completeness and Accuracy in Independent Validation
Natalia Kazimierczak1, Nora Sultani2, Szymon Krzykowski2
1Kazimierczak Clinic, Dworcowa 13/u6a, 85-009 Bydgoszcz, Poland.
Abstract:
Background/Objectives: Artificial intelligence (AI) tools for orthodontic diagnosis are increasingly used in clinical practice; however, there is limited evidence regarding their performance in CBCT-based assessments. In this study, we evaluated the diagnostic reliability of the Diagnocat platform for categorical orthodontic diagnoses obtained from CBCT examinations. Methods: Fifty-nine patients who underwent large-field CBCT (13 × 16 cm) and lateral cephalograms within 30 days were included, and CBCT scans were processed using Diagnocat (v1.0). The platform's categorical outputs-sagittal skeletal class, vertical facial pattern, overbite category, and Dental Angle class-were compared with manual cephalometric analyses performed by an experienced orthodontist (reference standard). Standard thresholds were used to convert reference continuous measurements into categorical variables. Missing or 'N/A' index test outputs were treated as diagnostic failures in accordance with STARD recommendations. Agreement was assessed via Cohen's kappa (κ), and the sensitivity, specificity, PPV, and NPV were calculated for angle classification. Results: The AI platform generated skeletal and vertical classifications in only 3/59 (5%) and 1/59 (1.7%) patients, respectively. Agreement was fair (κ = 0.324) for overbite categorization, and the Dental Angle class was provided for 34/59 (57.6%) patients. When "N/A" results were treated as diagnostic failures, the overall system usability was <10% for skeletal parameters. Conclusions: The platform demonstrated insufficient diagnostic reliability and failed to generate outputs for most patients. While the specificities for generated diagnoses were acceptable, the low data completeness rate renders the tool currently unsuitable for independent clinical decision-making.

