Related Experiment Video
Updated: Jun 4, 2026

Automated Joint Space Detection Improves Bone Segmentation Accuracy
Published on: November 28, 2025
Annotator reliability and probabilistic consensus for semantic segmentation in digital pathology
Laura Gálvez Jiménez1, Christine Decaestecker2
1Laboratory of Image Synthesis and Analysis, Université Libre de Bruxelles, Brussels, Belgium.
None:
In medical image segmentation and its semantic variant, annotations from multiple experts, when available, are used to generate consensus labels for training machine learning models. However, the annotator consistency is generally not captured in the consensus labels and can negatively impact the reliability of these models. In this study, we propose a novel approach based on the concept of self-consistency, which characterizes an expert's behavior by quantifying annotation certainty. The proposed method is model-agnostic and can be used as a universal preprocessing step for any segmentation backbone. To validate our approach, we apply it to semantic image segmentation in the context of prostate cancer grading, a domain known to be subject to both intra- and inter-expert variability. We use the PANDA dataset and generate synthetic experts to demonstrate that our approach provides valuable insights into intra-expert variability. Compared to other probabilistic approaches, such as soft and smooth labeling, our method improves the quality of the probabilistic consensus, thereby improving deep network training for semantic segmentation. We extend these findings to two real-world cancer datasets. The results highlight the potential of our approach to address challenges caused by annotation uncertainty in digital pathology. The code will be publicly available upon publication at https://github.com/lauragj95/SC_multi_expert.
