Related Experiment Video
Updated: May 4, 2026

Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
Published on: June 24, 2021
Combining DGE and RNA-sequencing data to identify new polyA+ non-coding transcripts in the human genome
Nicolas Philippe1, Elias Bou Samra, Anthony Boureux
1Transcriptomics, bioinformatics and myeloid leukaemia, INSERM, U1040, Institute for Research in Biotherapy, Montpellier F-34197, France, Université Montpellier 2, Montpellier, France, Institut de Biologie Computationnelle, Maison de la modélisation, Université Montpellier 2, France, LIRMM, MAB, CNRS UMR 5506, Université Montpellier 2, Montpellier, France and Genomic instability of pluripotent stem cells, INSERM, U1040, Institute for Research in Biotherapy, Montpellier F-34197, France.
Researchers identified ~34,000 novel transcribed regions using digital gene expression (DGE) and RNA-sequencing. This bioinformatics approach aids in discovering tissue-specific transcripts and understanding genome complexity.
Area of Science:
- Genomics
- Bioinformatics
- Transcriptomics
Background:
- Next-generation sequencing, particularly digital gene expression (DGE), enables massive parallel production of short reads for transcriptome analysis.
- DGE technologies offer a large dynamic range by generating short tag signatures for cell transcripts, which can be mapped to a reference genome.
- These tags can identify new transcribed regions, further characterized by RNA-sequencing (RNA-Seq) reads.
Purpose of the Study:
- To explore novel transcriptional regions and their biological features, including tissue expression and conservation.
- To integrate diverse data types, including DGE tags, RNA-Seq, tiling array expression data, and species comparison.
- To develop and apply a bioinformatics approach for comprehensive transcriptome analysis.
Main Methods:
- Analysis of a large digital gene expression dataset ('TranscriRef').
- Annotation of 750,000 uniquely mapped tags to the human genome using Ensembl.
- Integration of DGE tags, RNA-Seq, tiling array data, and species comparison to identify and categorize novel transcribed regions.
- Categorization of transcripts based on origin (protein-coding, antisense, intronic, intergenic) and overlap with non-coding transcripts.
Main Results:
- Identification of approximately 34,000 novel transcribed regions outside the boundaries of known protein-coding genes.
- Annotation of 750,000 unique tags mapped to the human genome.
- Characterization of novel transcripts originating from both DNA strands, including antisense, intronic, and intergenic regions.
- Demonstration of biological validation using sequencing data from human pluripotent stem cells.
Conclusions:
- The integrated bioinformatics approach successfully identified a significant number of novel transcribed regions.
- The method facilitates the discovery of tissue-specific candidate transcripts.
- This approach enhances the understanding of genome complexity and transcriptional landscape.
- The DigitagCT tool is available for broader application in transcriptome analysis.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Sanger Sequencing
Genomics
Genome Annotation and Assembly

