Related Experiment Video
Updated: Dec 28, 2025

05:07
Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
Published on: November 7, 2025
252
Variability in estimated gene expression among commonly used RNA-seq pipelines.
Sonali Arora1, Siobhan S Pattwell1, Eric C Holland2
1Division of Human Biology, Fred Hutchinson Cancer Research Center, Seattle, WA, 98109, USA.
Scientific Reports
|February 19, 2020
Summary
RNA-sequencing (RNA-seq) analysis shows most gene expression is consistent across pipelines. However, over 12% of genes show major differences, impacting disease biomarker discovery.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- RNA-sequencing (RNA-seq) is crucial for identifying disease biomarkers and therapeutic targets.
- Accurate mRNA abundance estimation is fundamental for downstream RNA-seq analyses.
- Current RNA-seq processing pipelines are assumed to provide reliable expression level estimates.
Purpose of the Study:
- To evaluate the consistency of gene expression abundance estimates across different RNA-seq processing pipelines.
- To identify protein-coding genes with significant expression discrepancies among pipelines.
- To assess the impact of pipeline variability on disease-associated gene identification.
Main Methods:
- Analysis of RNA-seq data from 6,690 human tumor and normal tissue samples.
- Application of five distinct, best-in-class RNA-seq processing pipelines to identical datasets.
- Comparison of mRNA abundance estimates and expression fold changes across pipelines for >19,000 protein-coding genes.
Main Results:
- Approximately 88% of protein-coding genes exhibited similar expression profiles across all tested pipelines.
- Over 12% of protein-coding genes showed more than four-fold differences in abundance estimates between pipelines.
- Discordant gene expression profiles were observed for many widely studied disease-associated genes.
- Variability in abundance estimates was influenced by diverse patterns of inter-pipeline discordance.
Conclusions:
- Significant pipeline-dependent variability exists in RNA-seq abundance estimates for a notable subset of genes.
- This variability can affect the reliability of disease biomarker and therapeutic target discovery.
- A community-wide effort is required to establish gold standards for estimating mRNA abundance for discordant genes.
- The identified list of discordantly evaluated genes serves as a resource for more robust marker discovery.
Related Concept Videos
RNA-seq
11.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.6K
Ribosome Profiling
4.0K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
4.0K

