Related Experiment Video
Updated: Jan 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations
Young Su Ko1, Jonathan Parkinson1, Wei Wang1,2
1Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA 92093-0359.
Integrating text data and fusing embeddings significantly improves protein language models (pLMs). A new algorithm efficiently combines embeddings, achieving state-of-the-art results in key biological tasks.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning in biology
Background:
- Protein language models (pLMs) utilize pretrained embeddings for transfer learning in biology.
- Standard pLM objectives may generate representations misaligned with downstream biological tasks.
- Increasing model size does not guarantee improved representation quality.
Purpose of the Study:
- To investigate strategies for enhancing protein language model representations.
- To improve the utility of pLMs for diverse biological applications.
- To address limitations of current pLM embedding strategies.
Main Methods:
- Integrating biological text annotations via contrastive learning to create text-integrated pLMs (tpLMs).
- Employing embedding fusion to combine representations from multiple pLMs.
- Developing and applying a greedy forward selection algorithm for efficient embedding subset identification.
Main Results:
- No single pLM or tpLM consistently outperformed others across all tested tasks.
- Fusion of tpLM embeddings enhanced performance on most biological tasks.
- The greedy forward selection algorithm efficiently identified near-optimal embedding subsets, achieving state-of-the-art results in homologous sequence recovery and protein-protein interaction prediction.
Conclusions:
- Embedding fusion is a practical and scalable method for improving protein representations.
- Text integration and embedding fusion offer complementary strategies for advancing pLM capabilities.
- The developed greedy algorithm overcomes computational bottlenecks associated with embedding fusion.
Related Concept Videos
Tagging and Fusion Proteins
Improving Translational Accuracy
Improving Translational Accuracy
Protein-Protein Interfaces
Protein-protein Interfaces
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
