Related Experiment Video
Updated: Oct 11, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Sparse Regression in Cancer Genomics: Comparing Variable Selection and Predictions in Real World Data
Robert J O'Shea1, Sophia Tsoka2, Gary Jr Cook1,3
1Department of Cancer Imaging, School of Biomedical Engineering and Imaging Sciences, King's College London, London, UK.
This study introduces a novel method to evaluate gene interaction models using real cancer genomics data. L0L2 penalisation excelled in structural selection, while L1L2 penalisation improved coefficient recovery, outperforming traditional cross-validation.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Evaluating gene interaction models in cancer genomics is difficult due to uncertain true distributions.
- Existing methods using synthetic or incomplete experimental data have limitations.
- A real-world data-driven approach is needed for robust model comparison.
Purpose of the Study:
- To develop and apply a novel benchmarking approach for genomic model inference algorithms using real-world data.
- To compare the performance of LASSO, elastic net, best-subset selection, L0L1, and L0L2 penalisation methods.
- To assess the efficacy of algorithmic preselection versus internal cross-validation for model selection.
Main Methods:
- Extracted five large genomic datasets (n=4000) from Gene Expression Omnibus.
- Trained 'gold-standard' regression models on data subspaces (n=4000, p=500).
- Trained penalised regression models on smaller samples (n=25, 75, 150) and validated against gold-standard models, assessing variable selection and prediction accuracy.
Main Results:
- L1L2 penalisation showed the highest cosine similarity for coefficient recovery.
- L0L2 penalisation explained the most variance in test responses and achieved the highest variable selection F1 score.
- Algorithmic preselection significantly outperformed internal cross-validation across all evaluated metrics.
Conclusions:
- The study presents a novel, data-driven approach for comparing model selection methods in cancer genomics.
- Benchmarking datasets are publicly available for future research.
- L0L2 penalisation is recommended for structural selection, and L1L2 penalisation for coefficient recovery in genomic data analysis.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Cancer Survival Analysis
Survival Tree
Building a Survival Tree
Constructing a...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Comparing the Survival Analysis of Two or More Groups
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as: