Related Experiment Video
Updated: Aug 16, 2025

07:15
Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11.1K
SHINE: protein language model-based pathogenicity prediction for short inframe insertion and deletion variants
Xiao Fan1,2, Hongbing Pan3, Alan Tian4
1Department of Pediatrics, Columbia University, New York, NY, USA.
Briefings in Bioinformatics
|December 28, 2022
Summary
We developed SHINE, a new tool for predicting the pathogenicity of short inframe insertion and deletion variants (indels). SHINE improves variant interpretation in genetic disease studies by leveraging protein language models.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate prediction of variant pathogenicity is crucial for genetic disease studies.
- Inframe indels pose interpretation challenges due to limited training data.
- Existing methods rely on manually encoded features, limiting predictive power.
Purpose of the Study:
- To develop a novel pathogenicity predictor for short inframe indels.
- To improve the interpretation of inframe indel variants in human diseases.
- To leverage deep learning and protein language models for enhanced variant prediction.
Main Methods:
- Developed SHINE (SHort Inframe iNsertion and dEletion) predictor.
- Utilized pretrained protein language models for latent representation of indels and protein context.
- Employed supervised machine learning models with curated training data from ClinVar and gnomAD.
Main Results:
- SHINE demonstrated superior prediction performance compared to existing methods.
- The predictor showed improved accuracy for both deletion and insertion variants.
- Performance was validated on two independent test datasets.
Conclusions:
- Unsupervised protein language models offer valuable insights into protein information.
- SHINE represents a significant advancement in variant interpretation for genetic analyses.
- The developed method enhances the prediction of inframe indel pathogenicity.
Related Concept Videos
Nonsense-mediated mRNA Decay
10.7K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.7K
Point and Frameshift Mutations
54
Point mutations are genetic alterations involving the change of a single nucleotide base pair in DNA. Depending on how the alteration affects protein synthesis, they can lead to various consequences.Point mutations fall into the following types:Silent mutations occur when a nucleotide change does not alter the amino acid sequence due to the redundancy of the genetic code. For instance, changing ACC to ACA still encodes threonine, leaving the protein function unaffected. This occurs because...
54
Leaky Scanning
5.2K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Mutations
84.0K
Overview
84.0K
Translation
15.1K
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Proteins are...
Translation Produces the Building Blocks of Life
Proteins are...
15.1K
Exon Recombination
3.7K
The evolution of new genes is critical for speciation. Exon recombination, also known as exon shuffling or domain shuffling, is an important means of new gene formation. It is observed across vertebrates, invertebrates, and in some plants such as potatoes and sunflowers. During exon recombination, exons from the same or different genes recombine and produce new exon-intron combinations, which might evolve into new genes.
Exon shuffling follows “splice frame rules.” Each exon...
Exon shuffling follows “splice frame rules.” Each exon...
3.7K

