Related Experiment Video
Updated: Jul 12, 2026

Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
Evaluation of Accuracy and Reproducibility of an AI-Based Cephalometric Landmark Detection Software: A Comparative
Afaf El Merouani-Drissi1, Hajar Ben-Mohimd2, Afaf Alassfar1
1Resident, Department of Orthodontics, Faculté de Médecine Dentaire, Université Mohammed V de Rabat.
Background:
Automated cephalometric landmark detection using artificial intelligence (AI) has emerged as a promising aid in orthodontic diagnosis and treatment planning. However, its clinical usefulness depends on both accuracy of detection and quality of the reference annotations that are used for evaluation. Objective: To assess the accuracy, precision and stability of an AI-based cephalometric landmark detection system (WebcephTm) using a rigorously validated ground truth.
Material And Methods:
One hundred lateral cephalograms were independently annotated by expert orthodontists. A custom-developed application was used to extract and standardize landmark coordinates in millimeters from radiographic images. Inter-observer agreement validated the ground truth. AI-annotated landmarks were compared with expert annotations using Euclidean distance measurements. Landmark-dependent variability was examined, and the stability of AI behavior across images was evaluated using the coefficient of variation (CV).
Results:
Inter-observer analysis demonstrated excellent agreement (ICC=1; mean discrepancy 1,53+- 0.70 mm), confirming the reliability of the reference. The AI system showed a mean localization error of 3,16 mm, with only 42,7% of detections within the 2-mm clinical threshold. Performance was strongly landmark-dependent. Soft-tissue landmarks showed higher localization errors than skeletal landmarks, although this difference was not statistically significant (Mann-Witney U test, p > 0,05). CV analysis confirmed that AI errors were anatomically driven rather than randomly distributed.
Conclusions:
AI-based cephalometric detection by WebcephTm is reliable for well-defined skeletal landmarks, while soft-tissue points remain challenging. The observed limitations appear to be primarily related to anatomical and radiographic complexity rather than to reference annotation inaccuracy or algorithmic instability.

