Related Experiment Videos
Improving functional annotation of non-synonomous SNPs with information theory
1Departments of Biopharmaceutical Sciences and Pharmaceutical Chemistry, University of California, San Francisco, San Francisco, CA 94143-2240, USA.
Summary
Selecting informative features is key for predicting the functional effects of non-synonymous single nucleotide polymorphisms (nsSNPs). A greedy algorithm identified a subset of features that improved prediction accuracy and reduced false positives in computational classifiers.
Area of Science:
- Bioinformatics
- Computational Biology
- Genetics
Background:
- Automated functional annotation of non-synonymous single nucleotide polymorphisms (nsSNPs) relies on descriptive features of amino acid changes.
- Identifying optimal feature combinations is crucial for accurate computational prediction methods.
Purpose of the Study:
- To rank 32 descriptive features based on their mutual information with the functional effects of amino acid substitutions.
- To identify a subset of highly informative features using a greedy algorithm for improved nsSNP functional annotation.
Main Methods:
- Mutual information was calculated between 32 features and experimentally verified functional effects of amino acid substitutions.
- A greedy algorithm was employed to select a subset of the most informative features.
- A support vector machine (SVM) classifier was trained and cross-validated using selected feature sets.
Main Results:
- Feature ranking by mutual information correlated with SVM classification accuracy.
- A reduced feature set (two solvent accessibility features and one evolutionary feature) achieved comparable accuracy with 6% fewer false positives than a 32-feature set.
- The selected features outperformed a comprehensive set including physiochemical, electrostatic, flexibility, and binding interaction properties.
Conclusions:
- The proposed feature selection method is effective and simple to implement for computational nsSNP annotation.
- A small subset of informative features can yield high predictive accuracy, reducing computational complexity and false positive rates.
- Solvent accessibility and evolutionary conservation are critical features for predicting the functional impact of amino acid substitutions.