Related Experiment Video
Updated: Jun 19, 2026

DeepOmicsAE: Representing Signaling Modules in Alzheimer's Disease with Deep Learning Analysis of Proteomics, Metabolomics, and Clinical Data
Published on: December 15, 2023
Learning High-Resolution Protein Embeddings from Multimodal Data via Self-Supervised Integration
Yong-Jia Liang1,2,3, Qian-Yi Wang1,2,3, Qian Zhou1,2,3
1School of Biomedical Engineering, Southern Medical University, Guangzhou 510515, China.
Abstract:
A massive volume of multimodal protein data such as amino acid sequences, structures, gene ontology (GO) annotations, and microscope images has been accumulated, but the experimentally validated function-related annotations of proteins remain scarce. Accurate protein representation as low-dimensional vectors is a prerequisite for employing machine learning, which in turn provides a promising paradigm for large-scale protein annotation. Although many studies in recent years have attempted to learn protein representations using deep learning, most of them rely on unimodal data like sequences or structures, ignoring the inherently multimodal nature of proteins. To address this, we present self-SSGI, a multimodal self-supervised method that integrates sequence, structure, GO annotation, and image data to learn high-resolution protein embeddings. The method first designs a joint masked reconstruction strategy to extract amino acid-level features from sequences and structures, and then, integrates GO annotations and protein images by contrastive learning to obtain protein-level features. Subsequently, the multilevel features are fused through a cross-attention-based multimodal fusion module to produce a unified embedding for each protein. Trained on 96,862 proteins, the embeddings learned by self-SSGI were applied to downstream tasks including protein subcellular localization, molecular function prediction, and protein-protein interaction inference. Experimental results demonstrate that self-SSGI efficiently integrates multiple modalities and enhances protein representation, leading to performance that surpasses state-of-the-art methods across multiple protein annotation tasks, including on external data sets. This work provides a useful protein representation tool to support further computational research in bioinformatics.