Related Experiment Video
Updated: Apr 18, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
A tri-modal contrastive learning framework for protein representation learning
Li Zhang1, Han Guo1, Leah Schaffer2
1Department of Electrical and Computer Engineering, University of California, San Diego (UC San Diego), La Jolla, CA 92093, USA.
Protein foundation models now integrate sequence, 3D structure, and text data for enhanced representations. This multimodal approach, ProteinAligner, improves predictions of protein functions and properties.
Area of Science:
- Biochemistry
- Computational Biology
- Structural Biology
Background:
- Protein foundation models, particularly language models, excel at learning representations from amino acid sequences using self-supervised learning on large datasets.
- These sequence-based representations are effective for predicting protein functions and properties.
- Current models often neglect crucial data like 3D structures and scientific literature, limiting their scope.
Purpose of the Study:
- To develop a multimodal pretraining framework that integrates protein sequences, 3D structures, and literature text.
- To address the limitations of existing models in modality coverage and training strategies.
- To capture richer and more holistic protein representations by leveraging complementary data sources.
Main Methods:
- Proposed a multimodal pretraining framework integrating three modalities: protein sequences, 3D structures, and literature text.
- Utilized protein sequences as the anchor modality.
- Employed contrastive learning to align structural and textual modalities with the sequence modality.
Main Results:
- The developed framework, ProteinAligner, captures more comprehensive protein representations.
- ProteinAligner demonstrated superior performance across a diverse range of downstream tasks compared to existing state-of-the-art foundation models.
- The model showed significant improvements in predicting protein functions and properties.
Conclusions:
- Integrating multiple modalities (sequences, structures, text) enhances protein representation learning.
- The proposed multimodal framework effectively captures holistic protein information.
- ProteinAligner represents a significant advancement in foundation models for protein science, improving predictive accuracy for biological functions and properties.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conservation of Protein Domains
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein and Protein Structures
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks

