Related Experiment Videos
Synthetic-to-Clinical Ensemble Learning for Volumetric Breast Tumor Segmentation in Digital Breast Tomosynthesis
Cristina Alfaro1,2, Gabriel Guerra3,4, Claudia Prieto4,5
1Department of Medical Technology, Facultad de Ciencias de la Salud, Universidad de Tarapacá, Arica 1000000, Chile.
Abstract:
Background/Objectives: Digital breast tomosynthesis (DBT) provides quasi-three-dimensional breast imaging, but volumetric segmentation model development is constrained by limited expert-annotated clinical data. This proof-of-concept study evaluated the feasibility of synthetic-to-clinical breast tumor segmentation using synthetic model development, limited clinical fine-tuning, and heterogeneous ensemble learning, and aimed to establish a reproducible low-resource framework for volumetric DBT segmentation. Methods: Three architectures-3D U-Net, nnU-Net, and Attention U-Net-were evaluated using publicly available synthetic and clinical DBT datasets. Models initialized from the Mixed-size synthetic configuration were fine-tuned using 10 clinical development cases with five-fold cross-validation; 10 additional clinical cases were reserved for independent testing. Ensemble weights and threshold were evaluated from development-set out-of-fold predictions. Performance was assessed using Dice, Intersection-over-Union (IoU), precision, and recall, with paired non-parametric comparisons on the clinical test cohort. Results: nnU-Net achieved the highest mean Dice on the Large Tumor (0.864) and Mixed-size (0.841) synthetic test sets, while performance was lower in the Small Tumor configuration (0.560). On the independent clinical cohort, the ensemble achieved the highest mean Dice (0.518) and recall (0.630), compared with nnU-Net (Dice = 0.482), Attention U-Net (0.372), and 3D U-Net (0.367). After Holm correction, ensemble Dice was significantly higher than Attention U-Net and 3D U-Net, but not nnU-Net. Conclusions: Synthetic DBT data provided a useful source-domain foundation for volumetric segmentation under constrained annotation conditions, although a synthetic-to-clinical gap remained. Heterogeneous ensemble integration achieved the highest mean clinical Dice without demonstrating superiority over fine-tuned nnU-Net. Larger and more diverse clinical cohorts are required to establish generalizability.