Related Experiment Video
Updated: Jun 3, 2025

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
634
Best holdout assessment is sufficient for cancer transcriptomic model selection
Jake Crawford1, Maria Chikina2, Casey S Greene3,4
1Genomics and Computational Biology Graduate Group, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA.
Patterns (New York, N.Y.)
|January 8, 2025
Summary
Simpler genomic models are not always better. This study found that smaller or more regularized gene signatures do not necessarily generalize better across datasets or cancer types. Predictive models should be chosen based on performance on held-out data.
Area of Science:
- Genomics
- Statistical modeling
- Machine learning in biology
Background:
- Statistical modeling guidelines often favor simpler models for genomics due to potential benefits in cost, interpretability, and generalization.
- The assumption that smaller gene signatures lead to better generalization across diverse biological contexts and datasets is widely held but not always empirically validated.
Purpose of the Study:
- To empirically test the assumption that smaller gene signatures generalize better in mutation status prediction.
- To evaluate the impact of model selection strategies (cross-validation vs. cross-validation with regularization) on the generalization performance of genomic models.
Main Methods:
- Developed mutation status prediction models using both linear (LASSO logistic regression) and non-linear (neural networks) approaches.
- Assessed model generalization by testing performance across different datasets (cell lines to human tumors) and biological contexts (holding out entire cancer types).
- Compared model selection based solely on cross-validation performance versus a combination of cross-validation and regularization strength.
Main Results:
- The study did not find evidence that more regularized or smaller gene signatures consistently generalized better.
- This finding was consistent across both cross-dataset and cross-cancer type generalization problems.
- Both linear and non-linear modeling approaches yielded similar results regarding the generalization of regularized signatures.
Conclusions:
- The principle that simpler or more regularized models inherently generalize better in genomics may not hold true in all scenarios.
- When aiming for generalizable predictive models in genomics, prioritizing models with the best performance on held-out data or via cross-validation is recommended.
- Model size and regularization strength should not be the primary criteria for selecting generalizable predictive models over empirical performance.

