Related Experiment Video
Updated: Jan 28, 2026

Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
Efficient Implementation of Penalized Regression for Genetic Risk Prediction
Florian Privé1, Hugues Aschard2, Michael G B Blum1
1Laboratoire TIMC-IMAG, UMR 5525, University of Grenoble Alpes, CNRS, 38700 La Tronche, France florian.prive@univ-grenoble-alpes.fr michael.blum@univ-grenoble-alpes.fr.
Penalized logistic regression (PLR) offers improved prediction for Polygenic Risk Scores (PRS) over traditional methods. This efficient approach scales to large biobank data, enhancing genetic risk prediction for diseases and traits.
Area of Science:
- Genetics and Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- Polygenic Risk Scores (PRS) are crucial for identifying genetic predisposition to diseases.
- The common Clumping+Thresholding (C+T) method for PRS derivation uses univariate GWAS summary statistics.
- Jointly estimating SNP effects may significantly enhance PRS predictive performance compared to C+T.
Purpose of the Study:
- To develop and evaluate an efficient penalized logistic regression (PLR) method for computing PRS using individual-level data.
- To implement penalized linear regression for quantitative traits.
- To compare the predictive performance and scalability of PLR against C+T and random forests.
Main Methods:
- Developed an efficient penalized logistic regression (PLR) implementation for joint SNP effect estimation on large individual-level datasets.
- Included automatic hyper-parameter selection within the PLR implementation.
- Applied penalized linear regression for quantitative traits and compared PLR, C+T, and random forests on simulated and real data.
Main Results:
- PLR achieved equal or superior predictive performance compared to C+T across various scenarios, demonstrating scalability to biobank-scale data.
- In simulations, PLR improved AUC from 83% to 92.5%; in celiac disease, AUC increased from 82.5% to 89%.
- PLR predicted height (correlation ~65% vs ~55%) and breast cancer risk more accurately than C+T in large UK Biobank cohorts.
Conclusions:
- Penalized regression is a feasible and relevant method for PRS computation with large individual-level datasets.
- The efficient R package bigstatsr enables practical application of PLR for enhanced genetic risk prediction.
- PLR offers significant improvements in predictive power for both disease and quantitative traits.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
08:42Design and Implementation of an fMRI Study Examining Thought Suppression in Young Women with, and At-risk, for Depression
Published on: May 19, 2015
Related Concept Videos
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Correlation and Regression
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Microsoft Excel: Regression Analysis
To perform regression...
Predicting Molecular Geometry