Related Experiment Video
Updated: Aug 15, 2025

Author Spotlight: Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
Artificial Intelligence for Detecting Cephalometric Landmarks: A Systematic Review and Meta-analysis.
Germana de Queiroz Tavares Borges Mesquita1, Walbert A Vieira2, Maria Tereza Campos Vidigal3
1Postgraduate Program in Dentistry, School of Dentistry, São Leopoldo Mandic, Campinas, São Paulo, Brazil.
This systematic review and meta-analysis evaluated how well computer vision tools identify specific anatomical points on dental X-rays compared to traditional human marking. While these automated systems show potential, current evidence is limited by high variability in existing research and a general lack of high-quality, low-bias studies.
Area of Science:
- Orthodontics research within artificial intelligence diagnostics
- Medical imaging informatics and digital dentistry
Background:
Digital diagnostics in orthodontics currently face a significant knowledge gap regarding the reliability of automated systems. Prior research has shown that while computer vision offers potential, the field lacks a unified standard for evaluating performance. No prior work had resolved the inconsistencies in how different studies report the precision of these automated tools. That uncertainty drove the need for a comprehensive synthesis of existing data. Researchers have struggled to compare automated performance against traditional human-led annotation due to high heterogeneity in published reports. This gap motivated a rigorous assessment of current diagnostic capabilities in dental imaging. Existing literature remains fragmented, making it difficult for clinicians to adopt these technologies with full confidence. Establishing a baseline for performance is necessary to move these tools from experimental settings into routine clinical practice.
Purpose Of The Study:
The aim of this study was to evaluate the use of artificial intelligence for detecting cephalometric landmarks in digital imaging examinations. Researchers sought to address the lack of consensus regarding the accuracy of these automated tools in orthodontic practice. The team compared automated detection results directly against traditional manual annotation methods to determine relative precision. This investigation was motivated by the high heterogeneity observed in existing literature regarding diagnostic performance. By synthesizing data from multiple databases, the authors intended to clarify the current state of automated dental diagnostics. The study addresses the urgent need for a standardized assessment of these emerging technological advances. Understanding the reliability of these systems is essential for their potential integration into routine clinical workflows. This review provides a comprehensive overview of the current evidence base to guide future research and clinical application.
Main Methods:
Review approach involved a systematic search across nine distinct electronic databases to identify relevant literature. Two independent reviewers performed the selection process to ensure consistency and minimize subjective errors. Data extraction followed a standardized protocol to capture performance metrics from all eligible publications. The team assessed the risk of bias for every included study using the QUADAS-2 framework. Statistical synthesis utilized random-effects meta-analysis to calculate agreement rates at specific spatial thresholds. Researchers compared automated outputs directly against manual annotation to determine precision and accuracy. The team calculated 95% confidence intervals to provide a range for the observed agreement and divergence values. This structured methodology aimed to mitigate the impact of high heterogeneity found within the existing body of dental imaging research.
Main Results:
Key findings from the literature indicate that automated systems achieved agreement rates of 79% at a 2 mm threshold. At a 3 mm threshold, the agreement rate improved to 90% when compared to manual marking. The analysis revealed a mean divergence of 2.05 mm between the two diagnostic methods. The menton landmark demonstrated the most consistent performance with a standardized mean difference of 1.17. High heterogeneity was present across the studies, with I-squared values reaching 99% for agreement rates. Only three of the 40 included studies met the criteria for low risk of bias across all evaluated domains. These results suggest that while automated detection is feasible, its current precision varies widely across different clinical settings. The data highlight a significant gap between experimental performance and the requirements for reliable clinical diagnostic tools.
Conclusions:
Synthesis and implications suggest that automated detection of anatomical points remains a promising area for future dental diagnostics. The authors note that current evidence supporting these systems is characterized by very low certainty. Researchers emphasize that future investigations must prioritize testing the robustness of these algorithms across diverse patient populations. The findings indicate that while some landmarks show high agreement, overall performance varies significantly between different studies. Synthesis and implications highlight that the current lack of high-quality, low-bias research limits widespread clinical implementation. Authors propose that standardizing evaluation metrics will be necessary to improve the validity of future diagnostic comparisons. The team concludes that while automated methods are evolving, they do not yet replace the need for rigorous validation against human standards. Future work should focus on addressing the high heterogeneity identified in existing diagnostic performance reports.
Frequently Asked Questions
The researchers propose that automated systems achieve agreement rates of 79% at a 2 mm threshold and 90% at a 3 mm threshold. This performance is compared against manual annotation, which serves as the reference standard for identifying specific anatomical points in dental images.
The authors utilized the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool to evaluate the risk of bias. This instrument allows researchers to systematically categorize the methodological quality of included studies, identifying those with low versus high potential for systematic error.
The authors state that the menton landmark exhibited the lowest divergence between automated and manual methods, with a standardized mean difference of 1.17. This specific anatomical point appears easier for algorithms to identify consistently compared to other cephalometric markers.
The researchers performed a random-effects meta-analysis to synthesize data from 40 studies. This statistical approach accounts for the high heterogeneity (I2=99%) observed across the collected literature, providing a more reliable estimate of overall diagnostic performance.
The authors measured mean divergence to quantify the distance between automated and manual markings. They reported a mean divergence of 2.05 mm, which helps clinicians understand the typical spatial error associated with current computer vision applications in orthodontics.
The researchers propose that future studies must focus on testing the strength and validity of these tools in different patient samples. They suggest that current findings are limited by very low certainty of evidence, necessitating more rigorous, standardized testing protocols.

