Related Experiment Video
Updated: Aug 3, 2025

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Known sequence features explain half of all human gene ends.
Aleksei Shkurin1,2, Sara E Pour1,2, Timothy R Hughes1,2
1Department of Molecular Genetics, University of Toronto, Toronto, ON M5S 1A8, Canada.
The five core sequence features accurately identify most human gene ends defined by cleavage and polyadenylation (CPA) sites. Additional RNA-binding proteins and motifs play a minimal role in distinguishing functional from cryptic CPA sites.
Area of Science:
- Molecular Biology
- Genomics
- Bioinformatics
Background:
- Cleavage and polyadenylation (CPA) sites are crucial for defining eukaryotic gene termination.
- Key sequence elements, including UGUA, polyadenylation signal (PAS), U-rich, CA/UA, and downstream GU-rich elements (DSEs), are known to be involved in CPA.
- The sufficiency of these core elements in delineating functional CPA sites and the role of other factors remain unclear.
Purpose of the Study:
- To investigate whether the five primary sequence features are sufficient to delineate cleavage and polyadenylation (CPA) sites.
- To assess the contribution of individual sequence features to CPA site determination using discriminative models.
- To evaluate the impact of additional sequences and RNA-binding proteins (RBPs) on distinguishing functional from cryptic CPA sites.
Main Methods:
- Utilized standard discriminative models to analyze the contributions of individual sequence features to CPA.
- Developed models incorporating the five primary CPA sequence features.
- Assessed the performance boost from U1-hybridizing sequences and known RBP RNA binding motifs.
Main Results:
- Models using only the five primary CPA sequence features correctly identified constitutive CPA sites at the ends of coding genes for 59% of human genes.
- U1-hybridizing sequences offered a minor improvement in prediction accuracy.
- Incorporating all known RBP RNA binding motifs increased the prediction accuracy to only 61%, indicating a minimal role for these factors in distinguishing real from cryptic sites.
Conclusions:
- The five established sequence features are highly effective in predicting human gene ends, accounting for the majority of constitutive cleavage and polyadenylation (CPA) sites.
- Additional factors, including U1-hybridizing sequences and known RNA-binding protein motifs, contribute minimally to the accurate delineation of functional CPA sites.
- The core CPA machinery's sequence elements are largely sufficient for distinguishing functional gene ends from cryptic sites.
More Related Videos
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Organization of Genes
Replication in Eukaryotes
Many Proteins Orchestrate Replication at the Origin
Eukaryotic replication follows many of the same...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Exon Recombination
Exon shuffling follows “splice frame rules.” Each exon...
RACE - Rapid Amplification of cDNA Ends

