Related Experiment Video
Updated: Aug 26, 2025

06:41
In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
13.8K
E-SNPs&GO: embedding of protein sequence and function improves the annotation of human pathogenic variants
Matteo Manfredi1, Castrense Savojardo1, Pier Luigi Martelli1
1Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna, Bologna 40126, Italy.
Bioinformatics (Oxford, England)
|October 13, 2022
Summary
E-SNPs&GO predicts disease-related protein variations using AI. This method efficiently annotates large datasets of human single amino acid variants, aiding precision medicine.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Precision Medicine
Background:
- Massive DNA sequencing generates numerous human single-nucleotide polymorphisms (SNPs) in protein-coding regions.
- Distinguishing harmful protein variations from neutral ones is a key challenge in precision medicine.
- Artificial intelligence (AI) offers novel approaches for protein sequence encoding, bypassing traditional evolutionary database searches.
Purpose of the Study:
- To develop a novel computational method, E-SNPs&GO, for predicting disease-relatedness of protein variations.
- To leverage advanced protein language models and embedding techniques for efficient variant annotation.
- To provide a scalable and accurate tool for analyzing large-scale protein variant datasets.
Main Methods:
- E-SNPs&GO utilizes protein language models and embedding techniques for input encoding.
- The method encodes both protein sequences and Gene Ontology (GO) functional annotations.
- A large dataset of 101,146 human protein single amino acid variants across 13,661 proteins was used for training.
Main Results:
- The model was trained on a comprehensive dataset derived from public resources.
- On a blind test set of 10,266 variants, E-SNPs&GO achieved a Matthews Correlation Coefficient (MCC) of 0.72.
- The performance of E-SNPs&GO is comparable to existing state-of-the-art methods for variant pathogenicity prediction.
Conclusions:
- E-SNPs&GO is proposed as an efficient and accurate tool for large-scale annotation of protein variant datasets.
- The method demonstrates the potential of AI-driven approaches in identifying disease-associated genetic variations.
- The developed webserver and datasets are publicly available for broader research use.
Related Concept Videos
Genome Annotation and Assembly
19.2K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.2K
Signal Sequences and Sorting Receptors
5.5K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.5K
Conserved Binding Sites
4.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.3K
Conservation of Protein Domains Over Different Proteins
11.1K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.1K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Mutations
84.1K
Overview
84.1K

