Related Experiment Video
Updated: Aug 5, 2026

Evaluation of Colorectal Cancer Risk and Prevalence by Stool DNA Integrity Detection
Published on: June 8, 2020
Prediction of microsatellite instability in colorectal cancer based on tissue phenotypes inferred from pathological
Zhiwu Wang1, Yankun Liu2, Wei Xiong1
1Hebei Key Laboratory of Molecular Oncology, Tangshan 063000, China; Tangshan Key Laboratory of Cancer Prevention and Therapy, Tangshan 063000, China; Department of Chemoradiotherapy, Tangshan People's Hospital, Tangshan 063000, China.
Abstract:
The application of Multiple Instance Learning (MIL) for classifying Whole Slide Images (WSIs) has gained extensive use in recent years, primarily due to the high cost and time consumption associated with pixel-level annotation of WSIs, which is challenging to accomplish. The advancements in MIL for WSIs have predominantly concentrated on two fronts: the development of superior feature extractors (for instance, utilizing self-supervised learning for training feature extractors) and the formulation of enhanced instance aggregation strategies. Regrettably, the majority of the most advanced approaches have neglected phenotypic variances among instances when employing attention mechanisms. To capitalize on the disparities between instance tissues, we have introduced a phenotypic self-distillation approach to MIL. Our framework is composed of three components: i) a self-supervised feature extractor based on contrastive learning and a phenotype extractor pre-trained on the Kather100K dataset, which automatically provides 9-class tissue phenotype labels (e.g., tumor epithelium, stroma, lymphocytes) without requiring manual annotation, ii) the incorporation of a self-distillation loss between the features of instances and their phenotypes to augment the informational content of both perspectives, and iii) the aggregation of MIL instances for the final MSI prediction. The efficacy of this framework was evaluated on two datasets: the TCGA-CRC dataset was used for training and internal testing with a fixed 70%/30% split, while the Tangshan People's Hospital cohort served as an independent external validation set. On the TCGA-CRC dataset (n = 360; 65 MSI-H, 295 MSS), our model achieved an AUC of 0.8846 and an accuracy of 0.84, using a fixed 70%/30% train-test split. On the Tangshan People's Hospital dataset (n = 472; 56 MSI-H, 426 MSS), the model attained an AUC of 0.7258 and an accuracy of 0.70.

