Related Experiment Video
Updated: Jul 2, 2025

Author Spotlight: Unveiling Transmembrane Protein Family-Related Markers in Gastric Cancer and Implications for Targeted Therapies
Published on: September 15, 2023
Performance analysis of data resampling on class imbalance and classification techniques on multi-omics data for
Yuting Yang1, Golrokh Mirzaei2
1Department of Computer Science and Engineering, The Ohio State University, Columbus, Ohio, United States of America.
This study developed computational models using multi-omics data to accurately classify cancer types. Machine learning, particularly with Stochastic Gradient Descent, achieved over 99% accuracy in predicting liver and breast cancers.
Area of Science:
- Computational biology
- Genomics
- Machine learning in oncology
Background:
- Cancer remains a major global health challenge.
- Early detection and accurate classification are crucial for effective treatment and improved patient outcomes.
- Multi-omics data integration offers a powerful approach for understanding cancer complexity.
Purpose of the Study:
- To develop and evaluate computational models for classifying normal versus tumor samples.
- To integrate RNA sequencing, copy number variation (CNV), and DNA methylation data for enhanced cancer classification.
- To compare various machine learning algorithms and techniques for addressing class imbalance in cancer datasets.
Main Methods:
- Utilized The Cancer Genome Atlas (TCGA) dataset for liver cancer, breast cancer, and colon adenocarcinoma.
- Developed joint analysis models integrating RNA seq, CNV, and DNA methylation data.
- Evaluated 18 machine learning methods using AUC, precision, recall, and F-measure, and compared five class imbalance techniques, with Synthetic Minority Oversampling Technique (SMOTE) showing superior performance.
Main Results:
- The model using Stochastic Gradient Descent (SGD) with Support Vector Machine (SVM) achieved over 99% accuracy and AUC >= 0.999 for liver and breast cancer.
- For colon adenocarcinoma, both SGD and Sequential Minimal Optimization (SMO) achieved 100% accuracy and 1.000 for AUC, precision, recall, and F-measure.
- SMOTE was identified as the most effective technique for handling class imbalance in the cancer datasets.
Conclusions:
- Integrated multi-omics data and machine learning provide highly accurate methods for cancer sample classification.
- SGD and SMO are effective algorithms for tumor classification, demonstrating excellent performance across different cancer types.
- Advanced computational approaches are vital for improving cancer detection and aiding in the fight against cancer.
More Related Videos
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Related Concept Videos
Cancer Survival Analysis
Genomics