Related Experiment Video
Updated: May 16, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Aggregating residue-level protein language model embeddings with optimal transport
Navid NaderiAlizadeh1, Rohit Singh1,2
1Department of Biostatistics and Bioinformatics, Duke University, Durham, NC 27705, United States.
Protein language models (PLMs) generate variable-length embeddings, posing challenges for downstream tasks. Our novel optimal transport method creates fixed-length protein representations, outperforming traditional pooling methods.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Protein language models (PLMs) generate per-residue embeddings from protein sequences.
- Variable-length outputs from PLMs hinder protein-level prediction tasks requiring uniform inputs.
- Average pooling is a common but potentially suboptimal method for summarizing PLM outputs.
Purpose of the Study:
- To develop a novel method for converting variable-length PLM embeddings into fixed-length representations.
- To address the challenge of variable sequence lengths in protein representation learning.
- To improve the performance of protein-level prediction tasks using PLMs.
Main Methods:
- Utilized optimal transport theory to aggregate per-token PLM outputs.
- Conceptualized token embeddings as samples from a probability distribution.
- Employed sliced-Wasserstein distances to create fixed-length, protein-level embeddings.
Main Results:
- The proposed optimal transport method generates fixed-length embeddings independent of protein length.
- This method demonstrated superior performance compared to average pooling across various downstream tasks.
- Smaller PLMs using this method achieved performance comparable to larger PLMs using average pooling, especially for longer sequences.
Conclusions:
- Optimal transport provides an effective strategy for generating fixed-length protein embeddings from PLMs.
- The method enhances the utility of PLMs for protein-level predictions, particularly for long sequences.
- This approach enables more efficient use of smaller PLMs while maintaining high performance.
More Related Videos
Related Concept Videos
Cotranslational Protein Translocation
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
Translocation of Proteins into the Mitochondria
Sorting of outer membrane proteins:
Mitochondrial outer membrane proteins are of two types: the transmembrane, beta-barrel porins, and the membrane-anchored, alpha-helical proteins. Beta-barrel porin precursors are translocated by the TOM complex and inserted into the outer mitochondrial membrane by the SAM complex. In contrast,...
Transport Across the Golgi
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

