Related Experiment Video
Updated: Jul 2, 2025

Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
The frequency of pathogenic variation in the All of Us cohort reveals ancestry-driven disparities
Eric Venner1, Karynne Patterson2, Divya Kalra3
1Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA. venner@bcm.edu.
Insights
This study reveals disparities in genomic data, showing higher pathogenic variant rates in European ancestry groups compared to African and Latino/Admixed American groups. These findings highlight the need for inclusive data to advance precision medicine.
Area of Science:
- Genomics
- Precision Medicine
- Population Health
Background:
- Disparities in clinical genomic interpretation data are a known issue, yet empirical evidence is limited.
- The All of Us Research Program is a large-scale initiative collecting diverse genomic and health data from over a million participants.
- Understanding variant frequencies across diverse populations is crucial for equitable healthcare.
Purpose of the Study:
- To examine pathogenic and likely pathogenic variants within the All of Us Research Program cohort.
- To identify and quantify disparities in variant frequencies across different ancestral groups.
- To inform targeted precision medicine strategies by revealing data biases.
Main Methods:
- Analysis of whole-genome sequencing data from the All of Us Research Program cohort.
- Identification and categorization of pathogenic and likely pathogenic variants.
- Comparison of variant frequencies across European, African, and Latino/Admixed American ancestry subgroups.
- Cross-referencing findings with the gnomAD database for validation and discrepancy analysis.
Main Results:
- The European ancestry subgroup exhibited the highest rate of pathogenic variation (2.26%).
- Lower rates were observed in the African ancestry group (1.62%) and Latino/Admixed American ancestry group (1.32%).
- Pathogenic variants were most commonly found in genes associated with Breast/Ovarian Cancer and Hypercholesterolemia.
- Variant frequencies largely aligned with gnomAD data, with some exceptions noted and resolved.
Conclusions:
- Observed differences in pathogenic variant frequencies between ancestral groups suggest ascertainment biases in current knowledge.
- Some deviations may indicate actual differences in disease prevalence across populations.
- This research provides critical insights into genomic data disparities, paving the way for more equitable precision medicine.
Abstract:
Disparities in data underlying clinical genomic interpretation is an acknowledged problem, but there is a paucity of data demonstrating it. The All of Us Research Program is collecting data including whole-genome sequences, health records, and surveys for at least a million participants with diverse ancestry and access to healthcare, representing one of the largest biomedical research repositories of its kind. Here, we examine pathogenic and likely pathogenic variants that were identified in the All of Us cohort. The European ancestry subgroup showed the highest overall rate of pathogenic variation, with 2.26% of participants having a pathogenic variant. Other ancestry groups had lower rates of pathogenic variation, including 1.62% for the African ancestry group and 1.32% in the Latino/Admixed American ancestry group. Pathogenic variants were most frequently observed in genes related to Breast/Ovarian Cancer or Hypercholesterolemia. Variant frequencies in many genes were consistent with the data from the public gnomAD database, with some notable exceptions resolved using gnomAD subsets. Differences in pathogenic variant frequency observed between ancestral groups generally indicate biases of ascertainment of knowledge about those variants, but some deviations may be indicative of differences in disease prevalence. This work will allow targeted precision medicine efforts at revealed disparities.
More Related Videos
Related Concept Videos
Genetic Variation
Genes exist in different versions called alleles,...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Mutation, Gene Flow, and Genetic Drift
Single Nucleotide Polymorphisms-SNPs
Human Genetics
The complex relationship between genetics and psychology is observable through common biological components such...

