Related Experiment Video
Updated: Sep 2, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method
Markus Bujotzek1,2, Dimitrios Bounias3, Stefan Denner3
1Division of Medical Image Computing, German Cancer Research Center, Im Neuenheimer Feld 280, Heidelberg, 69120, Germany. markus.bujotzek@dkfz-heidelberg.de.
Abstract:
While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections such as contour disagreement, missing or additional structures, and confused labels. Federated noisy label learning (FNLL) aims to mitigate these effects, yet remains underused in practice as existing evidence is largely based on synthetic noise, simplified settings, and limited real-world noisy evaluation. We address this gap by introducing a benchmark suite that combines diverse real-world noisy datasets, deployment-relevant client-noise scenarios, and label-noise-targeted evaluation to support systematic FNLL assessment and informed method selection. The suite combines curated real-world noisy medical image segmentation datasets from diverse sources with a comprehensive federated segmentation framework including various client-noise scenarios and noise-targeted evaluation. To demonstrate its capabilities, we compare representative FNLL methods across approaches, including noise-aware aggregation, robust personalization, label correction, and sample selection. In-depth data analysis shows that real-world segmentation label noise occurs both in isolation and in combination of characterized noise types. The benchmark shows that FedSelect performs strongest within noise-targeted evaluation, underlines FedAvg as a competitive baseline, and provides a practically informed decision guide for FNLL method selection based on label-noise type, client-noise scenario, and technical-cost considerations. The presented suite provides a realistic and discriminative basis for FNLL evaluation in medical image segmentation and establishes a reusable foundation for fair benchmarking, dataset-specific label-noise characterization, and future method development under realistic federated settings. Code is available at https://github.com/MIC-DKFZ/FedSegNoiseBench .
