Related Experiment Video
Updated: Sep 16, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
SPLENDID incorporates continuous genetic ancestry in biobank-scale data to improve polygenic risk prediction across
Tony Chen1,2, Haoyu Zhang3, Rahul Mazumder4
1Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
Abstract:
Polygenic risk scores are widely used in disease risk stratification, but their accuracy varies across different ancestries. Recent methods leverage multi-ancestry data to improve accuracy in under-represented populations but require the labeling of individuals by ancestry. This poses practical challenges, given that clinical decisions are typically not based on ancestry, and many individuals may not fit into a pre-specified ancestry group. Here we propose SPLENDID, a penalized regression framework for large-scale individual-level data that models genetic ancestry as a continuum to produce a single prediction model without any ancestry labels. In extensive simulations and analyses in the All of Us Research Program (n = 224,364) and UK Biobank (n = 340,140), we show that SPLENDID significantly improved prediction accuracy over existing methods, particularly for non-European and admixed ancestries. SPLENDID stands as a valuable tool for robust risk prediction across diverse populations, reduced health disparities in genetic research, and fairer clinical implementation.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Polygenic Traits
Polygenic Traits
Single Nucleotide Polymorphisms-SNPs
Genetic Variation
Genes exist in different versions called alleles, which...
Human Genetics
The complex relationship between genetics and psychology is observable through common biological components such...

