Related Experiment Videos
A quantitative assessment of the Hadoop framework for analyzing massively parallel DNA sequencing data
Alexey Siretskiy1, Tore Sundqvist1, Mikhail Voznesenskiy2
1Department of Information Technology, Uppsala University, P.O. Box 337, Uppsala, SE-75105 Sweden.
Gigascience
|June 6, 2015
Summary
The Hadoop platform offers a more efficient and economically viable solution for analyzing large next-generation sequencing datasets compared to traditional high-performance computing. This framework is poised to play a crucial role in future biological data analysis as data volumes continue to grow.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- High-throughput sequencing generates massive datasets, overwhelming traditional bioinformatics software on high-performance computing (HPC).
- The Hadoop platform offers distributed storage, processing, data locality, and fault tolerance for big data challenges.
Purpose of the Study:
- To quantitatively compare the efficiency of Hadoop versus conventional HPC for short read mapping and variant calling.
- To evaluate Hadoop's scalability and economic viability for next-generation sequencing (NGS) data analysis.
Main Methods:
- Compared Hadoop and HPC using ten datasets up to 100 gigabases with the Crossbow pipeline.
- Implemented an improved preprocessor for Hadoop and a graphical user interface (GUI) using the CloudGene platform.
- Measured computing hours and efficiency as a function of data size.
Main Results:
- Hadoop demonstrated greater efficiency for biologically relevant data sizes, outperforming HPC for both split and un-split data.
- Quantified the advantages of Hadoop's data locality for NGS, showing classical architectures with network-attached storage do not scale.
- The improved Hadoop pipeline scales better than the HPC implementation.
Conclusions:
- Hadoop is a scalable and economically viable option for current NGS data sizes.
- Hadoop is expected to become increasingly important for biological data analysis due to growing dataset sizes.