Related Experiment Video
Updated: Aug 27, 2025

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Organizing the bacterial annotation space with amino acid sequence embeddings
Susanna R Grigson1, Jody C McKerral2, James G Mitchell2
1Flinders Accelerator for Microbiome Exploration, College of Science and Engineering, Flinders University, Adelaide, South Australia, 5042, Australia. susie.grigson@flinders.edu.au.
Amino acid sequence embeddings offer a novel framework for protein function inference and ontology development. This approach helps organize protein data and identify proteins with unknown functions for experimental characterization.
Area of Science:
- Computational biology
- Bioinformatics
- Genomics
Background:
- Protein function inference is a significant challenge due to the growing number of protein discoveries and limited functional characterization.
- Current protein annotations rely on human-curated ontologies, which may not encompass all potential protein functions accurately.
- Advances in natural language processing and machine learning enable the embedding of amino acid sequences into vector spaces.
Purpose of the Study:
- To present amino acid sequence embeddings as a systematic framework for the study of protein ontologies.
- To explore the utility of embeddings in organizing and understanding protein function annotations.
- To investigate the potential of embeddings for identifying and characterizing proteins with unknown functions.
Main Methods:
- Utilizing sequence embeddings to analyze the structure of protein functional classes.
- Applying embedding techniques to bacterial carbohydrate metabolism data within the SEED annotation system.
- Embedding amino acid sequences of Bacillus proteins with unknown functions.
Main Results:
- The bacterial carbohydrate metabolism class, with 29 functional labels, was found to contain 48 distinct clusters of embedded sequences.
- Embedding unknown Bacillus amino acid sequences revealed distinct clusters, suggesting shared biological roles among these proteins.
- The study demonstrates a potential for sequence embeddings to reveal finer-grained functional relationships than existing labels.
Conclusions:
- Amino acid sequence embeddings represent a powerful tool for creating more robust protein ontologies and improving protein sequence data annotation.
- Embeddings can facilitate the clustering of proteins with unknown functions, aiding in the selection of candidates for experimental characterization.
- This approach offers a systematic way to enhance our understanding of protein function and biological roles.
Related Concept Videos
Genome Annotation and Assembly
Protein Organization
Amino acids
Tagging and Fusion Proteins
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

