Related Experiment Video
Updated: Jan 14, 2026

06:52
Discovery of Driver Genes in Colorectal HT29-derived Cancer Stem-Like Tumorspheres
Published on: July 22, 2020
6.9K
Machine learning approach to identify significant genes and classify cancer types from RNA-seq data
Sultana Akter1, Ridwan Olamilekan Adesola2, Shreya Basnet3
1College of Medicine and Life Sciences, Biomedical Sciences Concentrate Bioinformatics, University of Toledo, Ohio, USA.
Global Medical Genetics
|October 27, 2025
Summary
Machine learning accurately classifies cancer types using RNA-seq data. Support Vector Machines achieved 99.87% accuracy, offering efficient biomarker discovery for personalized cancer diagnostics.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Cancer is a major global health burden, causing millions of deaths annually.
- Current cancer identification methods are slow, costly, and require significant resources.
- There is a critical need for faster, more efficient cancer detection and classification techniques.
Purpose of the Study:
- To evaluate machine learning algorithms for cancer type classification using RNA-seq gene expression data.
- To identify statistically significant genes associated with different cancer types.
- To assess the efficiency and accuracy of various machine learning models in cancer genomics.
Main Methods:
- Utilized the PANCAN RNA-seq dataset from the UCI Machine Learning Repository.
- Assessed eight machine learning classifiers: Support Vector Machines, K-Nearest Neighbors, AdaBoost, Random Forest, Decision Tree, Quadratic Discriminant Analysis, Naïve Bayes, and Artificial Neural Networks.
- Validated model performance using a 70/30 train-test split and 5-fold cross-validation.
Main Results:
- The Support Vector Machine model demonstrated the highest classification accuracy, achieving 99.87% under 5-fold cross-validation.
- Identified statistically significant genes through RNA-seq data analysis.
- Compared the performance of eight different machine learning algorithms for cancer classification.
Conclusions:
- Machine learning, particularly Support Vector Machines, shows significant potential for accurate and efficient cancer classification from RNA-seq data.
- This approach can accelerate biomarker discovery and aid in developing personalized cancer diagnostics and treatments.
- The study highlights the utility of computational methods in advancing cancer research and clinical applications.
Related Concept Videos
RNA-seq
11.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.7K
lncRNA - Long Non-coding RNAs
9.8K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
9.8K

