The frequency of pathogenic variation in the All of Us cohort reveals ancestry-driven disparities

Eric Venner1, Karynne Patterson2, Divya Kalra3

  • 1Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, USA. venner@bcm.edu.

Communications Biology
|February 20, 2024
PubMed

Insights

This study reveals disparities in genomic data, showing higher pathogenic variant rates in European ancestry groups compared to African and Latino/Admixed American groups. These findings highlight the need for inclusive data to advance precision medicine.

Area of Science:

  • Genomics
  • Precision Medicine
  • Population Health

Background:

  • Disparities in clinical genomic interpretation data are a known issue, yet empirical evidence is limited.
  • The All of Us Research Program is a large-scale initiative collecting diverse genomic and health data from over a million participants.
  • Understanding variant frequencies across diverse populations is crucial for equitable healthcare.

Purpose of the Study:

  • To examine pathogenic and likely pathogenic variants within the All of Us Research Program cohort.
  • To identify and quantify disparities in variant frequencies across different ancestral groups.
  • To inform targeted precision medicine strategies by revealing data biases.

Main Methods:

  • Analysis of whole-genome sequencing data from the All of Us Research Program cohort.
  • Identification and categorization of pathogenic and likely pathogenic variants.
  • Comparison of variant frequencies across European, African, and Latino/Admixed American ancestry subgroups.
  • Cross-referencing findings with the gnomAD database for validation and discrepancy analysis.

Main Results:

  • The European ancestry subgroup exhibited the highest rate of pathogenic variation (2.26%).
  • Lower rates were observed in the African ancestry group (1.62%) and Latino/Admixed American ancestry group (1.32%).
  • Pathogenic variants were most commonly found in genes associated with Breast/Ovarian Cancer and Hypercholesterolemia.
  • Variant frequencies largely aligned with gnomAD data, with some exceptions noted and resolved.

Conclusions:

  • Observed differences in pathogenic variant frequencies between ancestral groups suggest ascertainment biases in current knowledge.
  • Some deviations may indicate actual differences in disease prevalence across populations.
  • This research provides critical insights into genomic data disparities, paving the way for more equitable precision medicine.

Related Concept Videos

Genetic Variation01:25

Genetic Variation

Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
284
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
13.4K
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Mutation, Gene Flow, and Genetic Drift01:09

Mutation, Gene Flow, and Genetic Drift

In a population that is not at Hardy-Weinberg equilibrium, the frequency of alleles changes over time. Therefore, any deviations from the five conditions of Hardy-Weinberg equilibrium can alter the genetic variation of a given population. Conditions that change the genetic variability of a population include mutations, natural selection, non-random mating, gene flow, and genetic drift (small population size).
58.4K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
Human Genetics01:28

Human Genetics

Human genetics provides a profound framework for understanding the interplay between genetic predispositions and human psychology. At the heart of this discipline lies the study of how genes influence physical traits, behaviors, and susceptibility to diseases. Each person carries a unique genetic code that subtly or significantly shapes their psychological and behavioral landscape.
The complex relationship between genetics and psychology is observable through common biological components such...
569