Related Experiment Video
Updated: Jun 2, 2026

03:37
Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Associations of Normalization and Regularization with Machine Learning Overfitting in Cross-dataset Classification of
Fei Deng1, Lanjing Zhang1,2,3,4
1Department of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, Piscataway, NJ, USA.
Summary
Normalization improves machine learning (ML) performance within datasets but increases overfitting. Regularized models like LASSO offer better cross-dataset performance, suggesting cautious use of normalization in omics data analysis.
Area of Science:
- Bioinformatics
- Machine Learning in Omics Data Analysis
- Cancer Genomics
Background:
- Normalization is crucial for standardizing omics data to enhance machine learning (ML) model performance.
- The association between normalization techniques and overfitting (i.e., poor performance on unseen data) in ML models for omics data remains unclear.
- Investigating the interplay of normalization and regularization is essential for robust ML applications in omics.
Purpose of the Study:
- To examine the association between normalization methods and overfitting in machine learning models applied to omics data.
- To evaluate the impact of regularization techniques on mitigating overfitting in cross-dataset omics analyses.
- To compare the performance of ML models with and without normalization across different datasets.
Main Methods:
- Utilized three paired transcriptomic and clinical cancer datasets (lung adenocarcinoma, melanoma, glioblastoma).
- Applied ANOVA-based gene selection, six normalization methods, and six ML models for patient death classification.
- Compared intra-dataset and cross-dataset performance using balanced accuracy (BA) and area under the curve (AUC).
Main Results:
- Normalization consistently improved intra-dataset performance across all datasets.
- Normalization showed minimal or slightly decreased cross-dataset performance, indicating increased overfitting.
- The Least Absolute Shrinkage and Selection Operator (LASSO) model, without normalization, demonstrated superior cross-dataset performance.
Conclusions:
- Normalization enhances intra-dataset ML performance but exacerbates overfitting in cancer omics data.
- Regularized models, such as LASSO, effectively mitigate overfitting and yield robust cross-dataset results.
- Cross-dataset evaluation and the use of regularized models are recommended for reliable ML in omics; normalization should be applied judiciously.