Related Experiment Video
Updated: Apr 9, 2026

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Factors affecting the accuracy of a class prediction model in gene expression data
Putri W Novianti1, Victor L Jong2,3, Kit C B Roes4
1Biostatistics & Research Support, Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, 3508, GA, Utrecht, The Netherlands. Novianti-3@umcutrecht.nl.
The number of differentially expressed genes and fold change significantly impact classification model accuracy in non-cancerous gene expression data. These factors, along with within-class correlation, explain most performance variations.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Class prediction models show variable performance in clinical gene expression datasets.
- Previous cancer-focused studies highlight dataset and classification function influence on model accuracy.
- Limited research exists on how gene expression data characteristics impact classifier performance.
Purpose of the Study:
- To empirically identify data characteristics affecting predictive accuracy of classification models.
- To investigate these factors specifically in non-cancerous gene expression datasets.
Main Methods:
- Downloaded datasets from 25 studies meeting inclusion criteria.
- Selected nine classification functions (discriminant analyses, Bayes classifiers, tree-based, regularization/shrinkage, nearest neighbors).
- Built and evaluated nine class prediction models per dataset, recording study and data characteristics (disease, sample size, gene expression features).
Main Results:
- The number of differentially expressed genes and average fold change significantly impacted classification model accuracy.
- These factors individually explained up to 72% and 57% of prediction accuracy variation.
- Multivariable analysis identified these two factors plus within-class correlation as key drivers, explaining 91.5% of between-study variation.
Conclusions:
- Number of differentially expressed genes, fold change, and within-class correlation significantly influence classification model accuracy in non-cancerous datasets.
- These data characteristics are crucial for understanding and improving model performance.
- Findings extend knowledge beyond cancer research regarding factors affecting predictive modeling in genomics.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Accuracy and Precision
Accuracy and Precision
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Chromatin Structure Regulates pre-mRNA Processing
The chromatin structure, especially...

