Related Experiment Video
Updated: Jan 14, 2026

07:34
A Simple Stimulatory Device for Evoking Point-like Tactile Stimuli: A Searchlight for LFP to Spike Transitions
Published on: March 25, 2014
10.3K
Sparse polygenic risk score inference with the spike-and-slab LASSO
Junyi Song1, Shadi Zabad1, Archer Yang2
1School of Computer Science, McGill University, Montréal, QC H3A 0G4, Canada.
Bioinformatics (Oxford, England)
|October 17, 2025
Summary
We developed SSLPRS, a new method for predicting disease risk using genetic data. It improves accuracy and variable selection, especially for sparse genetic architectures, offering a significant advancement in polygenic risk score methods.
Area of Science:
- Genetics
- Bioinformatics
- Statistical genetics
Background:
- Large-scale biobanks offer rich data for genetic studies of complex traits.
- High-dimensional genomic data presents challenges for disease risk prediction, especially with limited sample sizes.
- Existing Polygenic Risk Score (PRS) methods have limitations in scalability or coefficient shrinkage.
Purpose of the Study:
- To introduce SSLPRS, a novel PRS method utilizing the Spike-and-Slab LASSO (SSL) prior.
- To develop a scalable inference algorithm for PRS using Genome-Wide Association Studies (GWAS) summary statistics.
- To evaluate the performance of SSLPRS compared to existing methods.
Main Methods:
- Developed a coordinate-ascent inference algorithm for SSLPRS operating on GWAS summary statistics.
- Utilized the Spike-and-Slab LASSO (SSL) prior for PRS inference.
- Validated the method through simulations and analysis of UK Biobank quantitative phenotypes.
Main Results:
- SSLPRS demonstrates competitive prediction accuracy and superior variable selection performance, particularly in sparse genetic architectures.
- Achieved over 50% improvement in positive predictive value in simulations.
- Selected variants in real phenotype analyses are enriched for significant genomic annotations and show improved replication rates.
Conclusions:
- SSLPRS provides a robust and scalable approach for disease risk prediction from genotype data.
- The method bridges theoretical frameworks of sparse Bayesian priors and penalized regression.
- SSLPRS offers enhanced variable selection and prediction accuracy, advancing polygenic risk score methodologies.
Related Concept Videos
Polygenic Traits
68.9K
When more than one gene is responsible for a given phenotype, the trait is considered polygenic. Human height is a polygenic trait. Studies have uncovered hundreds of loci that influence height, and there are believed to be many more. Due to the high number of genes involved, as well as environmental and nutritional factors, height varies significantly within a given population. The distribution of height forms a bell-shaped curve, with relatively few individuals in the population at the...
68.9K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Multiple Allele Traits
37.9K
The Concept of Multiple Allelism
37.9K

