Related Experiment Video
Updated: Sep 3, 2025

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
872
MSPJ: Discovering potential biomarkers in small gene expression datasets via ensemble learning.
HuaChun Yin1,2,3, JingXin Tao2, Yuyang Peng1
1Department of Neurosurgery, Xinqiao Hospital, The Army Medical University, Chongqing 400037, China.
Computational and Structural Biotechnology Journal
|July 27, 2022
Summary
MSPJ, a novel machine learning method, enhances the identification of differentially expressed genes (DEGs) in small transcriptome datasets. This approach improves feature selection and stability, crucial for understanding complex disease mechanisms with limited samples.
Area of Science:
- Transcriptomics
- Bioinformatics
- Machine Learning
Background:
- Differentially expressed genes (DEGs) are crucial for understanding complex disease pathogenesis.
- Detecting DEGs is challenging in small gene expression datasets due to limitations like sample availability and funding.
- Existing methods often lack the power and stability required for small sample analyses.
Purpose of the Study:
- To introduce MSPJ, a new machine learning approach for robust DEG identification in small transcriptome datasets.
- To improve the power and stability of DEG detection when sample sizes are limited.
- To provide a reliable method for feature selection in transcriptomic analyses.
Main Methods:
- MSPJ utilizes an ensemble learning strategy combining improved multiple random sampling with meta-analysis, Support Vector Machine-Recursive Feature Elimination (SVM-RFE), and permutation testing.
- The method was evaluated using 94 simulated datasets and benchmarked against 10 classical methods on 165 real-world datasets.
- Performance was specifically assessed for small gene expression datasets, particularly those with fewer than 30 samples.
Main Results:
- MSPJ demonstrated superior performance in identifying DEGs across most small gene expression datasets compared to ten classical methods.
- The method showed particular effectiveness in datasets with sample sizes below 30.
- MSPJ provides robust feature selection, enhancing the reliability of DEG identification in low-sample scenarios.
Conclusions:
- MSPJ is an effective machine learning approach for robust DEG identification in small transcriptome datasets.
- The method addresses the limitations of small sample sizes in transcriptomic research.
- MSPJ is expected to advance research into the molecular mechanisms of complex diseases and phenotypes.
Keywords:
AUC, area under the ROC curve (AUC)DEGs, differentially expressed genesDifferentially expressed genesFDR, false positive rateFeature selectionGA, genetic algorithmGEO, Gene Expression OmnibusGO, gene ontologyMSPJ, the Joint method of Meta-analysis, SVM-RFE, and Permutation testMachine learningRF, random forestROC, receiver operating characteristicRandom samplingSAM, significance analysis of microarraysSMDs, standardized mean differencesSNR, signal noise ratioSVM-RFE, support vector machines-recursive feature eliminationSmall sample sizemRMR, minimum-redundancy-maximum-relevance
