Related Experiment Video
Updated: Mar 2, 2026

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.4K
Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning
Kai Wang1, Liuming Shi1, Xue Li1
1Key Laboratory of Advanced Design and Intelligent Computing, Ministry of Education, School of Software Engineering, Dalian University, Dalian 116622, China.
Journal of Biomedical Informatics
|February 28, 2026
Summary
This study introduces MViTGene, a new framework for predicting spatial gene expression from histological images. The model effectively aligns imaging and transcriptomic data, improving prediction accuracy for various gene types.
Area of Science:
- Computational Biology
- Genomics
- Medical Imaging
Background:
- Understanding tissue organization and molecular phenotypes relies on predicting spatial gene expression from histological images.
- Current methods struggle with single-model limitations and poor alignment between image and transcriptomic data.
Purpose of the Study:
- To develop a unified multimodal learning framework for integrating histological imaging and spatial transcriptomics.
- To enable effective cross-modal alignment through a shared latent representation space.
Main Methods:
- Utilized a ResNet50 convolutional stem and MobileViT Transformer backbone for histological image encoding.
- Employed linear-GELU-dropout transformation blocks to project modalities into a shared latent space.
- Implemented a contrastive learning objective for cross-modal alignment between image and spot embeddings.
Main Results:
- MViTGene demonstrated significantly higher prediction accuracy compared to existing methods on the human liver Visium dataset.
- Achieved improvements of 20%, 33%, and 12% in predicting marker, highly expressed, and highly variable genes, respectively.
- The model accurately captures the correspondence between tissue morphology and gene expression.
Conclusions:
- MViTGene offers a computational tool for high-throughput spatial gene expression prediction.
- The framework balances performance and interpretability, enabling more reliable biological interpretation.
- Advances the field of spatial transcriptomics analysis through multimodal integration.
Related Concept Videos
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K

