Related Experiment Video
Updated: Jul 24, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Impact of labelling inaccuracy and image noise on tooth segmentation in panoramic radiographs using federated,
Johan Andreas Balle Rubak1, Khuram Naveed1, Sanyam Jain1
1Department of Dentistry and Oral Health, Aarhus University, 8000 Aarhus, Denmark.
Objectives:
Federated learning (FL) may mitigate privacy constraints, heterogeneous data quality, and inconsistent labelling in dental artificial intelligence (AI). FL was compared with centralized (CL) and local learning (LL) for tooth segmentation in panoramic radiographs across multiple data corruption scenarios.
Methods:
An Attention U-Net was trained on 2066 radiographs from 6 institutions across 4 settings: baseline (unaltered data), label manipulation (dilated/missing annotations), image-quality manipulation (additive Gaussian noise), and exclusion of one faulty client with corrupted data. FL was implemented via the Flower AI framework. Per-client training and validation loss trajectories were monitored for anomaly detection and a set of metrics (Dice, IoU, HD, HD95, and ASSD) were evaluated on a hold-out test set. From these metrics significance results were reported through Wilcoxon signed-rank test. CL and LL served as comparators.
Results:
Baseline: FL achieved a median Dice of 0.949 (ASSD: 1.332), slightly better than CL at 0.947 (ASSD: 1.371) and LL at 0.936-0.940 (ASSD: 1.519-1.698). Label manipulation: FL maintained the best median Dice score at 0.949 (ASSD: 1.465) versus CL's 0.942 (ASSD: 1.757) and LL's 0.930-0.940 (ASSD: 1.519-2.115). Similar performance was observed when 2 faulty clients were introduced. Image noise: FL led with a Dice at 0.949 (ASSD: 1.311); CL had a Dice of 0.948 (ASSD: 1.361); LL ranged from 0.932 to 0.940 (ASSD: 1.519-1.774). Similar performance was observed when 2 faulty clients were introduced, with CL performing slightly better than FL. Faulty client exclusion: FL showed a Dice of 0.948 (ASSD: 1.331) better than CL's 0.946 (ASSD: 1.393). Loss curve monitoring reliably flagged the corrupted site.
Conclusions:
FL matches or exceeds CL and outperforms LL across corruption scenarios while preserving privacy. Per-client loss trajectories provide an effective anomaly-detection mechanism and support FL as a practical, privacy-preserving approach for scalable clinical AI development.
Related Concept Videos
Confocal Fluorescence Microscopy
Three-Dimensional Microscopy in Microbiology

