Related Experiment Video
Updated: Aug 8, 2025

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Which data subset should be augmented for deep learning? a simulation study using urothelial cell carcinoma
Yusra A Ameen1, Dalia M Badary2, Ahmad Elbadry I Abonnoor3
1Department of Computer Science, Faculty of Computers and Information, Assiut University, Asyut, Egypt. yusra.amin@aun.edu.eg.
Data augmentation in digital histopathology is crucial but not standardized. Augmenting both the test set and the combined training/validation set before splitting yields optimal deep learning model performance.
Area of Science:
- Digital pathology
- Machine learning
- Deep learning
Background:
- Deep learning for digital histopathology is limited by scarce annotated data.
- Data augmentation methods are not standardized, impacting model development.
- This study systematically compares 11 augmentation strategies.
Purpose of the Study:
- To explore the impact of data augmentation timing and dataset subset application on deep learning model performance in digital histopathology.
- To identify optimal data augmentation strategies for improving model accuracy and reliability.
Main Methods:
- Compared 11 data augmentation strategies, varying timing and dataset subsets (training, validation, test).
- Utilized hematoxylin-and-eosin-stained urinary bladder slides, classifying images into inflammation, urothelial cell carcinoma, or invalid categories.
- Fine-tuned four pre-trained convolutional neural networks (Inception-v3, ResNet-101, GoogLeNet, SqueezeNet) for binary classification.
- Evaluated model performance using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve.
Main Results:
- Augmentation applied after test-set separation but before training/validation split yielded the best testing performance, though with optimistic validation accuracy due to information leakage.
- Augmentation before test-set separation also led to optimistic results.
- Augmenting the test set provided more accurate evaluation metrics with reduced uncertainty.
- Inception-v3 demonstrated the best overall testing performance among the evaluated models.
Conclusions:
- For digital histopathology, data augmentation should encompass the test set (post-allocation) and the combined training/validation set (pre-split).
- This approach balances performance optimization with reliable evaluation.
- Further research is needed to generalize these findings across different datasets and tasks.
More Related Videos
09:28Patient-derived Orthotopic Xenograft Models for Human Urothelial Cell Carcinoma and Colorectal Cancer Tumor Growth and Spontaneous Metastasis
Published on: May 12, 2019
13:01Industrialized, Artificial Intelligence-guided Laser Microdissection for Microscaled Proteomic Analysis of the Tumor Microenvironment
Published on: June 3, 2022