Related Experiment Video
Updated: Jul 1, 2025

08:35
Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
Published on: June 24, 2021
5.6K
How tool combinations in different pipeline versions affect the outcome in RNA-seq analysis
Louisa Wessels Perelo1, Gisela Gabernet1, Daniel Straub1
1Quantitative Biology Center (QBiC), University of Tübingen, Otfried-Müller-Str. 37, 72076 Tübingen, Baden-Württemberg, 72076, Germany.
NAR Genomics and Bioinformatics
|March 8, 2024
Summary
Comparing RNA-seq analysis tools is crucial for consistent results. Different nf-core/rnaseq pipeline settings showed high gene classification overlap, but biases in low-concentration genes and gene isoforms were observed.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Data analysis tools, including RNA sequencing (RNA-seq) pipelines, are frequently updated.
- Changes in these tools can impact the comparability of results over time.
- Ensuring consistency in long-term transcriptomic studies is essential for reliable data interpretation.
Purpose of the Study:
- To evaluate the impact of different workflow options within the nf-core/rnaseq pipeline on data comparability.
- To compare the performance of various bioinformatics tools and pipeline versions for RNA-seq analysis.
- To identify factors influencing differential gene expression analysis outcomes.
Main Methods:
- Five distinct nf-core/rnaseq pipeline configurations were tested: STAR+Salmon, STAR+RSEM, STAR+featureCounts, HISAT2+featureCounts, and pseudoaligner Salmon.
- These configurations were applied to three diverse biological datasets: human, Arabidopsis, and zebrafish, all including External RNA Control Consortium (ERCC) spike-ins.
- Comparative analyses focused on fold change ratios and differential expression of both endogenous genes and spike-ins.
Main Results:
- An 85% overlap in differential gene classification was observed across the tested pipeline settings.
- Genes identified with a bias were predominantly those with lower expression concentrations.
- The number of gene isoforms and exons significantly influenced analysis outcomes, with previous featureCounts versions showing higher sensitivity for single-isoform genes like ERCC spike-ins.
Conclusions:
- While high concordance exists, variations in pipeline settings can introduce biases, particularly for lowly expressed genes and genes with complex isoform structures.
- Maintaining consistency in pipeline versions is recommended for long-term RNA-seq studies to ensure data comparability.
- During transitions to new pipeline versions, running both old and new versions concurrently is advised to validate gene targeting consistency.
Related Concept Videos
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K

