Related Experiment Video
Updated: Sep 16, 2026

Building Up a High-throughput Screening Platform to Assess the Heterogeneity of HER2 Gene Amplification in Breast Cancers
Published on: December 5, 2017
Toward AI-Assisted Precision Diagnostics in Breast Cancer: Source-Group-Preserving Evaluation of Class-Imbalance
Md Saiful Arefin1, Mohammad Saiful Islam1, Md Serajun Nabi2
1School of Computer Science and Informatics, Department of Computer Science, Albukhary International University, Alor Setar 05200, Kedah, Malaysia.
Abstract:
Background/Objectives: Estrogen receptor immunohistochemistry (ER-IHC) exhibits heterogeneous staining and substantial class imbalance, complicating segmentation of ordered expression categories. We evaluated whether class weighting, minority-focused crop sampling, adaptive minority curriculum (AMC), and Focal-Tversky optimization improved a common ResUNet-DS backbone for segmenting background/non-target pixels and C1 ER-negative, C2 weak-positive, C3 moderate-positive, and C4 strong-positive foreground categories. Methods: The dataset comprised 220 paired 512×512 image-mask patches organized into 44 recovered five-patch source groups. A source-group-preserving nested five-fold design used outer folds of 45, 45, 45, 45, and 40 patches. Six controlled training conditions were compared, with four independently selected inner models ensembled for each condition and outer fold. Results: Unweighted random training achieved the best numerical mean for all five primary endpoints: foreground Dice (0.7479±0.0123), C2-C4 Dice (0.7098±0.0128), foreground IoU (0.6074±0.0147), foreground quadratically weighted kappa (0.9761±0.0040), and foreground-ordinal MAE (0.0446±0.0062). Weighted random training was the strongest weighted/minority-sensitive condition but did not exceed the unweighted reference. None of 25 paired outer-fold comparisons reached p<0.05; the minimum raw p-value was 0.0625, and all Holm-adjusted p-values were 0.3125. A 20,000-replicate paired source-group bootstrap preserved the same overall ordering. Conclusions: More complex imbalance-handling strategies did not improve aggregate segmentation over unweighted training. Post hoc analyses showed scope-dependent probability quality and error ranking. IHC4BC provided expression-ordering consistency but not direct external segmentation validation. The findings support preliminary source-group-preserving methodological evidence rather than patient-level, whole-slide, or clinical validity.