Related Experiment Video
Updated: May 8, 2026

Automated Dissection Protocol for Tumor Enrichment in Low Tumor Content Tissues
Published on: March 29, 2021
A multimodal knowledge-enhanced whole-slide pathology foundation model
Yingxue Xu1, Yihui Wang1, Fengtao Zhou1
1Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China.
None:
Computational pathology has advanced through foundation models, yet faces challenges in multimodal integration and capturing whole-slide context. Current approaches typically utilize either vision-only or image-caption data, overlooking distinct insights from pathology reports and gene expression profiles. Additionally, most models focus on patch-level analysis, failing to capture comprehensive whole-slide patterns. Here we present mSTAR (Multimodal Self-TAught PRetraining), the pathology foundation model that incorporates three modalities: pathology slides, expert-created reports, and gene expression data, within a unified framework. Our dataset includes 26,169 slide-level modality pairs across 32 cancer types, comprising over 116 million patch images. This approach injects multimodal whole-slide context into patch representations, expanding modeling from single to multiple modalities and from patch-level to slide-level analysis. Across oncological benchmark spanning 97 tasks, mSTAR outperforms previous state-of-the-art models, particularly in molecular prediction and multimodal tasks, revealing that multimodal integration yields greater improvements than simply expanding vision-only datasets.
More Related Videos
07:32Author Spotlight: Investigating Immune Cell Dynamics in the Tumor Microenvironment — Challenges and Innovations in Cancer Prognosis
Published on: April 12, 2024
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025