Knowledge-enhanced pretraining for vision-language pathology foundation model on cancer diagnosis
Xiao Zhou1, Luoyi Sun2, Dexuan He3
1Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China.
We developed KnowledgE-Enhanced Pathology (KEEP), a novel foundation model that integrates medical knowledge for improved cancer diagnosis. KEEP enhances vision-language models by incorporating disease knowledge, significantly boosting performance, especially for rare cancer subtypes.
Area of Science:
- Computational pathology and artificial intelligence in cancer diagnosis.
- The application of knowledge-enhanced pretraining within vision-language foundation models.
- Medical informatics and disease ontology integration for histopathology image analysis.
Background:
Computational pathology has undergone a significant transformation through the development of vision-language foundation models that process large-scale histopathology datasets for diagnostic purposes. Prior research has shown that these models typically rely on purely data-driven approaches to align visual imagery with textual descriptions extracted from medical reports. While these systems demonstrate impressive performance in common diagnostic tasks, they often lack an explicit framework for integrating structured medical knowledge into their learning processes. Existing architectures frequently struggle to capture the complex hierarchical relationships between different disease states and their associated morphological features during the pretraining phase. The reliance on unstructured image-text pairs limits the ability of these models to generalize across rare cancer subtypes or nuanced diagnostic categories that require expert-level reasoning. A lack of formal ontology during training prevents the model from understanding the biological context that connects disparate clinical findings across different organ systems. This absence of evidence motivated the development of a more sophisticated training paradigm that bridges the gap between raw data and expert medical taxonomies.
Purpose Of The Study:
This research introduces KnowledgE-Enhanced Pathology (KEEP) to systematically integrate structured disease knowledge into the pretraining phase of vision-language models for cancer diagnosis. The project seeks to reorganize millions of pathology image-text pairs into 143,000 semantically structured groups that reflect established disease ontology hierarchies. By aligning visual and textual representations within a hierarchical semantic space, the study aims to foster a deeper understanding of disease relationships and morphological patterns. The investigators designed this framework to improve the recognition of subtle morphological features that are often overlooked by conventional data-driven models during standard training. A secondary objective involves enhancing the diagnostic accuracy for rare cancer subtypes where training data is typically scarce and expert knowledge is most valuable. The team intended to demonstrate that knowledge-guided learning provides a more robust foundation for computational pathology than standard vision-language alignment techniques used in current models. Ultimately, the work aims to prove that incorporating 11,454 diseases into the training loop significantly improves model robustness and diagnostic precision across diverse clinical settings.
Main Methods:
The researchers developed the KEEP foundation model by leveraging a comprehensive disease knowledge graph containing 11,454 distinct diseases and 139,143 unique attributes. This extensive graph served as the structural backbone for organizing millions of pathology image-text pairs into 143,000 semantically structured groups for hierarchical alignment. The pretraining process utilized these groups to align visual features with textual descriptions according to specific disease ontology hierarchies defined within the knowledge graph. The team evaluated the model's performance across 18 public benchmarks comprising over 14,000 whole-slide images (WSIs) to ensure broad diagnostic coverage. To test generalizability in challenging scenarios, the study incorporated 4 institutional rare cancer datasets encompassing 926 individual cases for validation. The methodology focused on creating a hierarchical semantic space where the model could learn both broad diagnostic categories and fine-grained morphological details simultaneously. This approach allowed the system to map visual representations of cancer cells directly to the structured attributes defined within the knowledge graph for improved interpretability.
Main Results:
KEEP consistently outperformed existing vision-language foundation models across all 18 public benchmarks and the 4 institutional rare cancer datasets tested in the study. The model demonstrated substantial performance gains specifically for rare cancer subtypes, which often pose significant challenges for traditional Artificial Intelligence (AI) systems due to data scarcity. By utilizing the structured knowledge graph, the system achieved a more precise alignment between visual morphological patterns and their corresponding textual attributes during the evaluation phase. The results from the 14,000 whole-slide images indicated that hierarchical pretraining improves the model's ability to distinguish between closely related disease states with high accuracy. In the institutional rare cancer cohorts, the model maintained high diagnostic performance despite the limited number of available training examples per subtype compared to common cancers. The findings suggest that the integration of 139,143 attributes into the pretraining phase significantly enhances the semantic depth of the learned representations for pathology. These quantitative improvements were observed across a diverse range of cancer types, validating the effectiveness of the knowledge-enhanced approach for clinical applications.
Conclusions:
Knowledge-enhanced vision-language modeling represents a powerful paradigm shift for advancing the field of computational pathology and improving cancer diagnosis workflows. The successful integration of disease ontology hierarchies into AI pretraining provides a blueprint for more reliable and interpretable cancer diagnosis tools in the future. These findings indicate that structured medical knowledge is essential for overcoming the limitations of purely data-driven foundation models in complex medical domains. Future developments in pathology AI will likely depend on the ability to synthesize large-scale image data with expert-curated knowledge graphs for better performance. The researchers anticipate that this approach will facilitate the discovery of new morphological biomarkers across a wide range of oncological conditions and patient populations. The study establishes a new standard for how foundation models should be architected to serve the complex needs of clinical pathology and precision medicine. By bridging the gap between machine learning and medical expertise, KEEP offers a scalable solution for high-precision diagnostic support in modern oncology.
Frequently Asked Questions
The KEEP model utilizes a knowledge graph of 11,454 diseases to organize image-text pairs into 143,000 semantically structured groups. This hierarchy aligns visual morphological patterns with textual attributes, allowing the system to better understand complex disease relationships during the pretraining phase.
The researchers integrated a comprehensive knowledge graph encompassing 139,143 attributes to reorganize millions of pathology image-text pairs. This structured approach enabled the model to outperform existing foundation models across 18 public benchmarks containing over 14,000 whole-slide images.
The knowledge graph provided a disease ontology hierarchy that allowed the researchers to align visual and textual representations within a structured semantic space. This method enabled the KEEP model to capture morphological patterns and diagnostic relationships that purely data-driven models often fail to identify.
The study focused on improving diagnostic performance for rare cancer subtypes, which often lack sufficient training data in standard datasets. By using structured knowledge, KEEP achieved substantial gains across 4 institutional rare cancer datasets comprising 926 individual cases.
The study's authors propose that knowledge-enhanced vision-language modeling is a powerful paradigm for advancing computational pathology. They conclude that integrating structured medical knowledge into pretraining is essential for developing more accurate and robust AI-driven cancer diagnosis systems.
Related Concept Videos
Cancer Survival Analysis
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Cancer
