Genome-wide identification of dominant polyadenylation hexamers for use in variant classification

Henoke K Shiferaw1, Celine S Hong1, David N Cooper2

  • 1Center for Precision Health Research, National Human Genome Research Institute, National Institutes of Health, 50 South Drive, Bethesda, MD 20892, United States.

Human Molecular Genetics
|August 22, 2023
PubMed

Insights

This study defines functionally important polyadenylation signal hexamers genome-wide, identifying a resource to better investigate genetic variants linked to human diseases. These hexamer variants can now be systematically analyzed for clinical relevance.

Area of Science:

  • Molecular Biology
  • Genetics
  • Bioinformatics

Background:

  • Polyadenylation is crucial for mRNA stability and cytoplasmic export.
  • Polyadenylation signal hexamers are key to this process, but variants are under-recognized in Mendelian disorders.
  • Current hexamer definitions are limited, hindering clinical analysis.

Purpose of the Study:

  • To define functionally important polyadenylation signal hexamers genome-wide.
  • To create a resource for interrogating hexamer variants in research and clinical settings.
  • To establish criteria for classifying pathogenic hexamer variants.

Main Methods:

  • Identified predominant polyA sites (pPAS) and associated hexamers with >50% usage per gene.
  • Assessed population constraint and variant burden (CADD scores) for predominant hexamers.
  • Interrogated exome data for hexamer variants and analyzed transcriptome data (3' RNA-seq) for functional impact.

Main Results:

  • Defined 4532 predominant hexamers with significant population constraint and higher CADD scores.
  • Identified 65 variants in predominant hexamers in 1477 individuals.
  • Observed that variants in predominant hexamers led to alternative polyadenylation (38%) and elongated transcripts (12%).

Conclusions:

  • The identified predominant hexamers serve as a valuable resource for interrogating genomic data.
  • Variants in these hexamers are strong candidates for pathogenic variation in human diseases.
  • Proposed pathogenicity criteria will aid in classifying hexamer variants for clinical use.

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
13.6K
Histone Variants at the Centromere02:30

Histone Variants at the Centromere

Histone variants are the histone proteins with structural and sequence variations. These variants may be regarded as “mutant” forms that replace their canonical histone counterparts in the nucleosomes. Specific post-translational modifications on the histone variants enable further chromatin complexity and regulate tissue-specific gene expression. The most common histone variants are from histone H2A, H2B, and linker histone H1 families. However, several variants of histone H3...
4.4K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.3K
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K