Univariate-guided sparse regression for Biobank-scale high-dimensional omics data
Joshua Richland1, Tuomo Kiiskinen2, William Wang3
1Department of Statistics, Stanford University, Stanford, California, United States of America.
Abstract:
We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (uniLasso). UniLasso is a two-stage penalized regression procedure that leverages univariate coefficients and magnitudes to stabilize feature selection and produce sparse predictive models. Building on its theoretical and empirical advantages, we adapt uniLasso for application to the UK Biobank, a population-based repository comprising over one million genetic variants measured on hundreds of thousands of individuals from the United Kingdom. We further extend the framework to incorporate external summary statistics via uniLasso ES (external scores). These signals guide the regression toward variants with prior evidence of association by informing penalty weights and sign constraints. Both uniLasso ES and uniLasso ultimately fit multivariate models using individual-level target data; the external statistics guide, rather than replace, this fitting. Our results demonstrate that uniLasso attains predictive performance comparable to standard Lasso while selecting substantially fewer variants, yielding sparser and potentially more interpretable models. Moreover, it remains competitive with other PRS estimation methods, such as PRS-CS and lassosum2.
Related Concept Videos
Biostatistics: Overview
Discrete variables are...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Quantifying and Rejecting Outliers: The Grubbs Test
Statistical Software for Data Analysis and Clinical Trials

