Related Experiment Video
Updated: Jun 21, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.8K
Masked pre-training of transformers for histology image analysis
Shuai Jiang1, Liesbeth Hondelink1, Arief A Suriawinata2
1Department of Biomedical Data Science, Geisel School of Medicine at Dartmouth, Hanover, NH 03755, USA.
Journal of Pathology Informatics
|July 15, 2024
Summary
MaskHIT, a self-supervised vision transformer (ViT) model, effectively pre-trains on whole-slide images (WSIs) for digital pathology tasks. This approach enhances WSI analysis for cancer diagnosis and prediction, outperforming existing methods.
Area of Science:
- Digital pathology
- Computational pathology
- Artificial intelligence in medicine
Background:
- Whole-slide images (WSIs) are crucial in digital pathology for cancer diagnosis and prognosis.
- Vision transformer (ViT) models show promise for analyzing WSIs but face challenges due to large parameter counts and limited labeled data.
Purpose of the Study:
- To develop a self-supervised pre-training method for ViT models applied to WSIs.
- To improve the performance of ViT models on WSI-level tasks like survival prediction, cancer subtype classification, and grade prediction.
Main Methods:
- Proposed MaskHIT, a novel pretext task for self-supervised pre-training of transformer models on WSIs.
- Utilized contrastive loss for reconstructing masked patches based on transformer outputs.
- Pre-trained the MaskHIT model on over 7000 WSIs from The Cancer Genome Atlas (TCGA).
Main Results:
- Pre-training with MaskHIT enables context-aware WSI understanding and learning of histological features.
- Achieved superior performance compared to multiple instance learning and state-of-the-art transformer methods.
- Demonstrated a 3% and 2% improvement on survival prediction and cancer subtype classification tasks, respectively.
- Attention maps from MaskHIT align with pathologist annotations, identifying clinically relevant structures.
Conclusions:
- Self-supervised pre-training is essential for optimal ViT performance on WSI-level tasks.
- MaskHIT facilitates the extraction of representative histological features based on patch position and visual patterns.
- The model accurately identifies clinically relevant histological structures, aiding in cancer diagnosis and prognosis.

