Related Experiment Video
Updated: Nov 22, 2025

06:41
In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
14.0K
Predicting the Disease Risk of Protein Mutation Sequences With Pre-training Model
Kuan Li1,2, Yue Zhong3, Xuan Lin4
1School of Cyberspace Security, Dongguan University of Technology, Guangdong, China.
Frontiers in Genetics
|January 7, 2021
Summary
BertVS accurately predicts disease risk from missense mutations in tumor suppressor genes like BRCA1 and PTEN. This framework integrates protein sequence context and amino acid properties for improved variant classification.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Missense mutations in tumor suppressor genes (e.g., BRCA1, PTEN) can cause protein dysfunction and increase disease risk.
- Accurate identification of disease-associated missense mutations is crucial for risk assessment and potential therapeutic strategies.
Purpose of the Study:
- To develop a hybrid framework, BertVS, for predicting the disease risk associated with protein missense mutations.
- To leverage deep learning and biochemical properties for enhanced variant classification.
Main Methods:
- Pre-training BERT models on the Pfam database to capture protein sequence context.
- Integrating amino acid hydrophilic properties with sequence representations.
- Concatenating learned representations and classifying missense mutations using a classifier.
Main Results:
- The pre-trained BERT model achieved 0.984 accuracy on masked language model prediction.
- BertVS demonstrated superior performance on BRCA1 and PTEN datasets, achieving 0.920 AUROC and 0.915 AUPR.
- The framework effectively classified known clinical variants in the ClinVar dataset.
Conclusions:
- BertVS successfully learns functional information from protein sequences.
- The proposed method effectively predicts disease risk for missense variants, including those with uncertain clinical significance.
- BertVS offers a promising tool for variant interpretation in cancer genomics.
Related Concept Videos
Signal Sequences and Sorting Receptors
12.7K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
12.7K
Protein Families
16.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.3K
Conserved Binding Sites
4.8K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.8K
Mutations
91.9K
Overview
91.9K
Mutations
42.0K
Mutations are changes in the sequence of DNA. These changes can occur spontaneously or they can be induced by exposure to environmental factors. Mutations can be characterized in a number of different ways: whether and how they alter the amino acid sequence of the protein, whether they occur over a small or large area of DNA, and whether they occur in somatic cells or germline cells.
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
42.0K

