Related Experiment Video
Updated: Oct 21, 2025

03:08
Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
331
Landscape of the Dark Transcriptome Revealed Through Re-mining Massive RNA-Seq Data
Jing Li1,2,3, Urminder Singh2,3,4, Zebulun Arendsee2,3,4
1Genetics and Genomics Graduate Program, Iowa State University, Ames, IA, United States.
Frontiers in Genetics
|September 6, 2021
Summary
Researchers explored the yeast "dark transcriptome," finding many unannotated sequences (ORFs) are transcribed and likely code for proteins. This discovery opens new avenues for understanding gene function and evolution.
Area of Science:
- Genomics and Transcriptomics
- Molecular Biology
- Bioinformatics
Background:
- The "dark transcriptome" comprises transcribed sequences lacking gene annotation.
- Understanding these unannotated regions is crucial for a complete view of the genome's functional potential.
Purpose of the Study:
- To investigate the expression and potential protein-coding capacity of unannotated open reading frames (ORFs) in the Saccharomyces cerevisiae genome.
- To identify candidate protein-coding genes within the dark transcriptome.
- To provide a framework for exploring and reusing public RNA-Seq data for novel gene discovery.
Main Methods:
- Analyzed 3,457 RNA-Seq samples from Saccharomyces cerevisiae under diverse conditions.
- Evaluated expression of 6,692 annotated genes and 29,354 unannotated ORFs.
- Utilized phylostratigraphic analysis and Markov Chain Clustering for ORF classification and co-expression analysis.
- Integrated data into MetaOmGraph (MOG) for interactive visualization and analysis.
Main Results:
- Over 30% of highly transcribed ORFs demonstrated translation evidence.
- Phylostratigraphic analysis suggested most transcribed ORFs encode species-specific proteins ("orphan-ORFs").
- Hundreds of orphan-ORFs exhibited expression levels comparable to annotated genes.
- Markov Chain Clustering identified 2,468 orphan-ORFs within co-expression clusters.
Conclusions:
- A significant portion of the dark transcriptome in yeast likely represents unannotated protein-coding genes.
- These findings highlight the potential for novel gene discovery within previously overlooked genomic regions.
- The MetaOmGraph platform facilitates the exploration of large-scale RNA-Seq data, enabling hypothesis generation for experimental validation.
Related Concept Videos
RNA-seq
10.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.6K
Ribosome Profiling
3.7K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.7K

