Related Experiment Video
Updated: Feb 16, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Analyzing large scale genomic data on the cloud with Sparkhit
Liren Huang1,2,3, Jan Krüger1,2, Alexander Sczyrba1,2,3
1Faculty of Technology, Bielefeld University, Bielefeld 33615, Germany.
Motivation:
The increasing amount of next-generation sequencing data poses a fundamental challenge on large scale genomic analytics. Existing tools use different distributed computational platforms to scale-out bioinformatics workloads. However, the scalability of these tools is not efficient. Moreover, they have heavy run time overheads when pre-processing large amounts of data. To address these limitations, we have developed Sparkhit: a distributed bioinformatics framework built on top of the Apache Spark platform.
Results:
Sparkhit integrates a variety of analytical methods. It is implemented in the Spark extended MapReduce model. It runs 92-157 times faster than MetaSpark on metagenomic fragment recruitment and 18-32 times faster than Crossbow on data pre-processing. We analyzed 100 terabytes of data across four genomic projects in the cloud in 21 h, which includes the run times of cluster deployment and data downloading. Furthermore, our application on the entire Human Microbiome Project shotgun sequencing data was completed in 2 h, presenting an approach to easily associate large amounts of public datasets with reference data.
Availability And Implementation:
Sparkhit is freely available at: https://rhinempi.github.io/sparkhit/.
Contact:
asczyrba@cebitec.uni-bielefeld.de.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
Genomics
Evolutionary Relationships through Genome Comparisons
DNA Microarrays
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...

