Related Experiment Video
Updated: Oct 18, 2025

08:04
Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
671
Accurate Sequence-Based Prediction of Deleterious nsSNPs with Multiple Sequence Profiles and Putative Binding
Ruiyang Song1, Baixin Cao1, Zhenling Peng2
1School of Mathematical Sciences, Nankai University, Tianjin 300071, China.
Biomolecules
|September 28, 2021
Summary
Predicting disease-causing non-synonymous single nucleotide polymorphisms (nsSNPs) is crucial. Our new sequence-based predictor, DMBS, significantly improves the accuracy of identifying deleterious nsSNPs, aiding disease research.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Non-synonymous single nucleotide polymorphisms (nsSNPs) can cause pathogenic changes linked to human diseases.
- Accurate prediction of deleterious nsSNPs is essential for understanding disease mechanisms.
- Current nsSNP prediction tools offer modest performance, necessitating improved methods.
Purpose of the Study:
- To develop a novel, highly accurate sequence-based predictor for deleterious nsSNPs.
- To enhance the prediction of pathogenic nsSNPs by leveraging protein sequence conservation and functional site information.
- To provide a tool that can effectively guide experimental validation in high-throughput settings.
Main Methods:
- Developed DMBS, a sequence-based predictor utilizing improved conservation estimates from multiple sequence profiles (two databases, two alignment algorithms).
- Integrated putative functional/binding residue annotations from state-of-the-art sequence-based methods.
- Employed a random forests model for classification, empirically compared against five other machine-learning algorithms.
Main Results:
- DMBS achieved an Area Under the Curve (AUC) > 0.94 across four benchmark datasets, outperforming existing methods.
- DMBS demonstrated superior performance on specific datasets, achieving AUC = 0.97 for SNPdbe and AUC = 0.97 for ExoVar, compared to 0.70 and 0.88 for prior methods.
- On the independent HumVar dataset, DMBS significantly outperformed the state-of-the-art SNPdryad method.
Conclusions:
- DMBS offers a significant advancement in predicting deleterious nsSNPs with high accuracy.
- The method's reliance on enhanced conservation estimates and functional site predictions contributes to its superior performance.
- DMBS provides a valuable tool for efficiently prioritizing nsSNPs for experimental validation in disease research.
More Related Videos
Related Concept Videos
Conserved Binding Sites
4.7K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.7K
Single Nucleotide Polymorphisms-SNPs
16.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
16.8K
Signal Sequences and Sorting Receptors
10.7K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
10.7K

