Related Experiment Video
Updated: Jun 2, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Associations of Normalization and Regularization with Machine Learning Overfitting in Cross-dataset Classification of
Fei Deng1, Lanjing Zhang1,2,3,4
1Department of Chemical Biology, Ernest Mario School of Pharmacy, Rutgers University, Piscataway, NJ, USA.
Background And Objectives:
Normalization can standardize and improve machine learning (ML) performance on omics data. However, it is unclear whether normalization is associated with overfitting (i.e., worse cross-dataset performance than intra-dataset performance). Therefore, we aimed to examine associations of normalization and regularization with overfitting of ML on omics data.
Methods:
Using three paired transcriptomic and clinical datasets (lung adenocarcinoma: the Cancer Genome Atlas (TCGA)/Oncology Singapore; melanoma: TCGA/Dana-Farber Cancer Institute; glioblastoma: TCGA/Clinical Proteomic Tumor Analysis Consortium), we applied ANOVA-based gene selection methods, six normalization methods, and six ML models to classify cancer patients' deaths. Balanced accuracy (BA) and area under the curve (AUC) in intra- and cross-dataset settings were compared using inferential analyses.
Results:
Normalization consistently improved intra-dataset performance (median BA/AUC changes: 0.035-0.214/0.115-0.279) on all data, particularly with Z_Raw, but decreased or slightly increased cross-dataset performance (median BA/AUC changes: -0.029-0.079/0.029-0.064). Least Absolute Shrinkage and Selection Operator (LASSO) model without normalization consistently outperformed most of the ML models in cross-dataset testing across cancer types. ML models on all and molecular-alone data showed similar best performances.
Conclusions:
Normalization increases ML's intra-dataset performance and overfitting in three paired cancer transcriptomic and clinical datasets. Regularized models such as LASSO appear to mitigate overfitting and achieve robust cross-dataset performance. Therefore, cross-dataset evaluation and regularized models are recommended to assess and reduce overfitting, while normalization should be used cautiously. Adding clinical data seems to have little impact on ML models' performance. However, future work on other diseases and datasets is warranted.