Genome-wide identification of dominant polyadenylation hexamers for use in variant classification
Henoke K Shiferaw1, Celine S Hong1, David N Cooper2
1Center for Precision Health Research, National Human Genome Research Institute, National Institutes of Health, 50 South Drive, Bethesda, MD 20892, United States.
Abstract:
Polyadenylation is an essential process for the stabilization and export of mRNAs to the cytoplasm and the polyadenylation signal hexamer (herein referred to as hexamer) plays a key role in this process. Yet, only 14 Mendelian disorders have been associated with hexamer variants. This is likely an under-ascertainment as hexamers are not well defined and not routinely examined in molecular analysis. To facilitate the interrogation of putatively pathogenic hexamer variants, we set out to define functionally important hexamers genome-wide as a resource for research and clinical testing interrogation. We identified predominant polyA sites (herein referred to as pPAS) and putative predominant hexamers across protein coding genes (PAS usage >50% per gene). As a measure of the validity of these sites, the population constraint of 4532 predominant hexamers were measured. The predominant hexamers had fewer observed variants compared to non-predominant hexamers and trimer controls, and CADD scores for variants in these hexamers were significantly higher than controls. Exome data for 1477 individuals were interrogated for hexamer variants and transcriptome data were generated for 76 individuals with 65 variants in predominant hexamers. 3' RNA-seq data showed these variants resulted in alternate polyadenylation events (38%) and in elongated mRNA transcripts (12%). Our list of pPAS and predominant hexamers are available in the UCSC genome browser and on GitHub. We suggest this list of predominant hexamers can be used to interrogate exome and genome data. Variants in these predominant hexamers should be considered candidates for pathogenic variation in human disease, and to that end we suggest pathogenicity criteria for classifying hexamer variants.
Insights
This study defines functionally important polyadenylation signal hexamers genome-wide, identifying a resource to better investigate genetic variants linked to human diseases. These hexamer variants can now be systematically analyzed for clinical relevance.
Area of Science:
- Molecular Biology
- Genetics
- Bioinformatics
Background:
- Polyadenylation is crucial for mRNA stability and cytoplasmic export.
- Polyadenylation signal hexamers are key to this process, but variants are under-recognized in Mendelian disorders.
- Current hexamer definitions are limited, hindering clinical analysis.
Purpose of the Study:
- To define functionally important polyadenylation signal hexamers genome-wide.
- To create a resource for interrogating hexamer variants in research and clinical settings.
- To establish criteria for classifying pathogenic hexamer variants.
Main Methods:
- Identified predominant polyA sites (pPAS) and associated hexamers with >50% usage per gene.
- Assessed population constraint and variant burden (CADD scores) for predominant hexamers.
- Interrogated exome data for hexamer variants and analyzed transcriptome data (3' RNA-seq) for functional impact.
Main Results:
- Defined 4532 predominant hexamers with significant population constraint and higher CADD scores.
- Identified 65 variants in predominant hexamers in 1477 individuals.
- Observed that variants in predominant hexamers led to alternative polyadenylation (38%) and elongated transcripts (12%).
Conclusions:
- The identified predominant hexamers serve as a valuable resource for interrogating genomic data.
- Variants in these hexamers are strong candidates for pathogenic variation in human diseases.
- Proposed pathogenicity criteria will aid in classifying hexamer variants for clinical use.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Histone Variants at the Centromere
Single Nucleotide Polymorphisms-SNPs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...


