Rare coding and noncoding variants map 1,342 diseases and biomarkers in 490,549 whole genomes
Yuxin Yuan1, Yuanyuan Guan1, Yannuo Feng1
1School of Mathematics and Statistics and KLAS, Northeast Normal University, Changchun, Jilin, China.
Medrxiv : the Preprint Server for Health Sciences
|April 3, 2026
Summary
This study analyzed whole genome sequencing data from nearly half a million UK Biobank participants to map rare genetic variants associated with diseases and biomarkers. The findings reveal numerous novel noncoding variant associations, offering insights for drug discovery.
Area of Science:
- Genomics
- Human Genetics
- Statistical Genetics
Background:
- Rare genetic variants, particularly noncoding ones, significantly contribute to human trait heritability but are understudied in large biobanks.
- These variants are often biologically specific and less polygenic than common variants, necessitating specialized analytical approaches.
Purpose of the Study:
- To comprehensively assess the impact of rare coding and noncoding variants across a wide range of human phenotypes using whole genome sequencing data.
- To develop and apply a scalable framework for phenome-wide association studies of rare variants.
Main Methods:
- Analysis of whole genome sequencing (WGS) data from up to 490,549 UK Biobank participants.
- Development and application of STAARpipelinePheWAS, a scalable framework for WGS phenome-wide rare variant association analysis.
- Assessment of 1,342 phenotypes, including diseases, clinical biomarkers, and metabolomics traits.
Main Results:
- Identification of 49,121 genome-wide significant gene-trait pairs, highlighting the role of rare variants in human traits.
- Discovery of numerous associations involving noncoding rare variants, many previously undetected by exome or array-based studies.
- Enrichment of identified associations in drug targets and biologically relevant pathways, suggesting translational potential.
Conclusions:
- This study provides a comprehensive map of rare noncoding variant associations across disease and biomarker domains.
- The findings offer a foundational resource for rare variant discovery, functional interpretation, and advancing translational genomics.
- Public accessibility of results via an interactive portal facilitates further research and application.
Related Concept Videos
Genome-wide Association Studies-GWAS
16.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.7K
Pharmacogenomics: Identification of New Drug Targets
86
Advances in genomics have profoundly influenced drug discovery by increasing both the speed and accuracy of pharmaceutical development. Pharmacogenomics, which examines how genetic variation influences drug response, facilitates the identification of novel therapeutic targets and enables patient stratification for personalized treatment. These strategies contribute to improved drug efficacy, minimized adverse effects, and more efficient clinical trial design.Mapping genetic differences...
86
Comparing Copy Number Variations and SNPs
19.3K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.3K
Modern Molecular Taxonomy
837
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
837
Principles of Pharmacogenetics: Types of Genetic Variants
97
The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
97
Genomics
41.8K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
41.8K


