Survey of Protein Sequence Embedding Models
Chau Tran1, Siddharth Khadkikar2, Aleksey Porollo3,4,5
1Department of Computer Science, University of Cincinnati, Cincinnati, OH 45219, USA.
International Journal of Molecular Sciences
|February 25, 2023
Summary
Protein language models generate numerical embeddings for diverse protein sequences. This study evaluates models for yeast proteome analysis, gene ontology annotation, and human disease variant prediction, finding significant differences in pathogenic mutations.
Area of Science:
- Computational biology
- Bioinformatics
- Protein science
Background:
- Natural language processing (NLP) algorithms have led to protein language models (PLMs) that encode protein sequences into fixed-size numerical vectors (embeddings).
- These PLMs are crucial for analyzing diverse protein lengths and amino acid compositions in various biological contexts.
- Representative models include Esm, Esm1b, ProtT5, SeqVec, and their derivatives like GoPredSim and PLAST.
Purpose of the Study:
- To survey and evaluate representative protein embedding models for diverse computational biology tasks.
- To assess the performance of these models in embedding the Saccharomyces cerevisiae proteome.
- To investigate the utility of PLMs for gene ontology (GO) annotation, human disease variant analysis, antimicrobial resistance correlation, and fungal mating factor analysis.
Main Methods:
- Surveyed and applied established protein embedding models (Esm, Esm1b, ProtT5, SeqVec) and their derivatives.
- Embedded the Saccharomyces cerevisiae proteome using these models.
- Utilized embeddings for GO annotation of uncharacterized proteins.
- Analyzed embeddings of human protein variants in relation to disease status.
- Correlated embeddings of Escherichia coli beta-lactamase TEM-1 mutants with experimental antimicrobial resistance data.
- Examined embeddings of fungal mating factors.
Main Results:
- All surveyed models indicated that uncharacterized yeast proteins are typically short (<200 amino acids), less rich in aspartate and glutamate, and enriched in cysteine.
- High-confidence GO term annotation was achieved for less than half of the uncharacterized yeast proteins.
- A statistically significant difference was observed in cosine similarity scores between benign and pathogenic human protein mutations.
- Embeddings of TEM-1 mutants showed low to no correlation with experimentally measured minimal inhibitory concentrations (MICs).
Conclusions:
- Protein language models offer valuable insights into protein characteristics and functions across different organisms and contexts.
- PLMs demonstrate potential for identifying patterns in uncharacterized proteins and predicting disease associations, though with limitations.
- The correlation between protein embeddings and quantitative functional measures like MICs requires further investigation and model refinement.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
11.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.0K
Per-Unit Sequence Models
107
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
107
Protein Organization
6.7K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.7K
Conservation of Protein Domains
3.1K
3.1K
Protein and Protein Structures
10.6K
10.6K
Protein and Protein Structure
80.0K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
80.0K


