Related Experiment Video
Updated: Dec 30, 2025

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.6K
SparkRA: Enabling Big Data Scalability for the GATK RNA-seq Pipeline with Apache Spark
Zaid Al-Ars1, Saiyi Wang1, Hamid Mushtaq1
1Computer Engineering Lab, Delft University of Technology, Mekelweg 5, 2628 CD Delft, The Netherlands.
Genes
|January 18, 2020
Summary
SparkRA, a new pipeline, significantly speeds up RNA sequencing (RNA-seq) variant calling by leveraging Apache Spark. This scalable solution drastically reduces computational time for analyzing large RNA-seq datasets, making complex genomic analyses more accessible.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- RNA sequencing (RNA-seq) data is rapidly increasing, driving demand for efficient analysis techniques.
- Current RNA-seq variant calling pipelines, like GATK's, face computational bottlenecks limiting scalability for large datasets.
- Efficient analysis is crucial for applications such as genotype-phenotype relationship identification and result validation.
Purpose of the Study:
- To develop a scalable pipeline for RNA-seq variant calling using Apache Spark.
- To address the computational limitations of existing GATK RNA-seq best practices pipelines.
- To accelerate the analysis of large-scale RNA-seq data.
Main Methods:
- Developed SparkRA, an Apache Spark-based pipeline for GATK RNA-seq variant calling.
- Implemented parallel processing across multiple cores and compute nodes.
- Evaluated performance on a single node and a distributed cluster.
Main Results:
- SparkRA reduced computation time by approximately 4× on a single node (from >5 h to 1.3 h for 32 GB data).
- On a 16-node cluster, SparkRA achieved a 7.7× speedup compared to single-node performance.
- SparkRA demonstrated a 1.2× speed advantage over other scalable solutions with equivalent accuracy.
Conclusions:
- SparkRA offers a highly efficient and scalable solution for RNA-seq variant calling.
- The pipeline significantly overcomes computational limitations, enabling faster analysis of large genomic datasets.
- SparkRA enhances the accessibility and practicality of RNA-seq data analysis in various biological research areas.
More Related Videos
Related Concept Videos
RNA-seq
11.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.6K
Sanger Sequencing
772.4K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.4K
Maximum Size of Aggregate
469
The maximum size of aggregate is defined as the aperture of the sieve retaining 15 percent or more of the particles present in the aggregate sample. The aggregate's maximum size impacts the concrete's water requirement, workability, and strength. Larger aggregates reduce the surface area needing cement paste coverage, which can lower water needs, thereby allowing a decrease in the water-to-cement ratio when the desired workability and richness of the mix are to be maintained, which can...
469
Next-generation Sequencing
97.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.4K
RACE - Rapid Amplification of cDNA Ends
7.0K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
7.0K

