Related Experiment Video
Updated: Apr 23, 2026

06:52
Discovery of Driver Genes in Colorectal HT29-derived Cancer Stem-Like Tumorspheres
Published on: July 22, 2020
6.0K
Identification and analysis of driver missense mutations using rotation forest with feature selection
1Key Laboratory of Intelligent Computing and Signal Processing of Ministry of Education, Anhui University, Hefei, Anhui 230601, China ; School of Computer Science and Technology, Anhui University, Hefei, Anhui 230601, China.
Biomed Research International
|September 25, 2014
Summary
This study introduces DX, a novel feature selection method for identifying cancer driver mutations using machine learning. The method achieved high accuracy in predicting these critical genetic alterations.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying cancer-associated mutations, or driver mutations, is crucial for understanding cancer genome function, including oncogene activation and tumor suppressor gene inactivation.
- Existing supervised machine learning approaches for driver mutation prediction often lack clarity on the importance of specific features.
- This highlights a need for effective feature selection methods in cancer genomics research.
Purpose of the Study:
- To propose and evaluate a novel feature selection method, termed DX, for identifying important features in cancer driver mutation prediction.
- To assess the performance of the DX method using the rotation forest algorithm.
- To demonstrate the generalizability of the DX method across different cancer datasets.
Main Methods:
- A novel feature selection method (DX) was developed using a set of 126 candidate features.
- The rotation forest algorithm was employed for experimental validation.
- The method was trained and tested on datasets from the COSMIC and Swiss-Prot databases.
Main Results:
- The DX method, utilizing the 11 top-ranked features, achieved high prediction performance: 88.03% accuracy, 93.9% precision, and 81.35% recall.
- The selected features proved effective in identifying cancer-associated mutations.
- Comparative analysis on TP53, EGFR, and Cosmic2plus datasets confirmed the method's generality.
Conclusions:
- The DX feature selection method is effective for identifying critical features in cancer driver mutation prediction.
- The rotation forest algorithm combined with DX yields high predictive performance.
- The proposed method demonstrates broad applicability across various cancer-related datasets.
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
14.2K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.2K
Comparing Copy Number Variations and SNPs
11.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.4K

