Related Experiment Video
Updated: Sep 16, 2026

Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
Cross-Domain Generalization of Deep Learning Architectures for Cephalometric Landmark Detection: A Dual-Dataset and
1Department of Orthodontics, Faculty of Dentistry, Istanbul Health and Technology University, 34275 Istanbul, Türkiye.
Abstract:
Background/Objectives: Deep learning models for cephalometric landmark detection report near-ceiling accuracy on single benchmarks, yet most are trained and tested on the same dataset. Whether the best in-domain architecture remains best out-of-domain has not been systematically quantified. Methods: Four architecture families (heatmap CNN, two-stage cascade, coordinate-regression Vision Transformer, pretrained ResNet-50) were each trained on two independently sourced datasets-ISBI 2015 (400 images, one device) and Aariz (1000 images, seven devices)-and evaluated on both, over their 19 shared landmarks in millimeters. All axes used three seeds (mean ± SD): the cross-dataset matrix, leave-one-device-out shift, balanced joint training, a pretrained-versus-scratch ablation, and a landmark-level breakdown. Results: In-domain mean radial error (MRE) was 2.47-3.04 mm (Aariz) and 4.12-5.55 mm (ISBI); cross-dataset error rose steeply (Generalization Drop-the relative increase in error out-of-domain-155-426%). The most accurate model in-domain (a from-scratch heatmap CNN, 2.47 mm) showed the largest drop, and every pretrained estimate fell below every from-scratch estimate (family means 248% vs. 380%): in-domain ranking did not predict cross-domain ranking. Leave-one-device-out revealed reproducible device-specific shift (held-out MRE 1.89-7.94 mm). Inter-observer variability was 0.53 mm, so cross-domain errors were 15-44× the human band. Balanced joint training reduced the cross-domain gap for all four architectures (both domains ≈ 2.1-3.5 mm) without harming in-domain accuracy. Pretraining more than halved cross-domain error from Aariz to ISBI (7.73 vs. 16.36 mm) but not in reverse, supporting the mechanism in one direction rather than uniformly. A-point, B-point, Nasion, and Menton all exceeded the 2 mm clinical threshold out-of-domain, including points localized to sub-millimeter accuracy in-domain. Conclusions: Single-dataset accuracy substantially overstates clinical generalizability, and the best in-domain architecture is not the most transferable, so in-domain leaderboards are an unreliable basis for selecting a model for clinical deployment; balanced multi-source training recovers most of the loss across the architectures tested.