Related Experiment Videos
Meta-analysis approach as a gene selection method in class prediction: does it improve model performance? A case
Putri W Novianti1,2,3, Victor L Jong4,5, Kit C B Roes4
1Biostatistics & Research Support, Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, 3508, GA, Utrecht, The Netherlands. p.novianti@vumc.nl.
BMC Bioinformatics
|April 13, 2017
Summary
Gene expression meta-analysis improves predictive model performance by selecting relevant genes. This approach is most effective for detecting low fold changes and highly correlated genes in acute myeloid leukemia data.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Meta-analysis of gene expression data enhances effect estimate precision and statistical power.
- This study investigates meta-analysis as a gene selection strategy for predictive modeling.
Purpose of the Study:
- To evaluate the benefit of using meta-analysis for gene selection before predictive modeling.
- To compare meta-analysis-based gene selection with single-dataset approaches.
Main Methods:
- Trained classification models on single gene expression datasets and validated externally.
- Performed gene selection via meta-analysis on four datasets, then trained and validated predictive models on separate datasets.
- Generated synthetic datasets to evaluate factors influencing model performance.
Main Results:
- Meta-analysis gene selection improved classification model performance for some datasets compared to single-dataset methods.
- Performance gains were dataset-dependent; some showed no major improvement.
- Fold change and pairwise correlation of differentially expressed genes significantly impacted the performance difference between meta-analysis and individual-classification models.
- Meta-analysis was more effective with low fold change and high pairwise correlation.
Conclusions:
- Gene selection via meta-analysis can enhance predictive model performance on gene expression data.
- The effectiveness depends on data characteristics like fold change and gene correlation.