Related Experiment Video
Updated: May 16, 2026

08:03
Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
READSCAN: a fast and scalable pathogen discovery program with accurate genome relative abundance estimation
Raeece Naeem1, Mamoon Rashid, Arnab Pain
1Pathogen Genomics Laboratory, Computational Bioscience Research Center, King Abdullah University of Science and Technology (KAUST), Thuwal-23955-6900, Kingdom of Saudi Arabia. raeece.naeem@gmail.com
Bioinformatics (Oxford, England)
|November 30, 2012
Summary
READSCAN efficiently identifies non-host pathogen sequences and estimates their abundance in large sequencing datasets. This scalable parallel program accurately classifies sequences, enabling rapid analysis of high-throughput data.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- High-throughput sequencing generates vast datasets requiring efficient analysis tools.
- Identifying non-host sequences, potentially of pathogen origin, is crucial for disease surveillance and research.
- Estimating the relative abundance of pathogen genomes aids in understanding infection dynamics.
Purpose of the Study:
- To introduce READSCAN, a scalable parallel program for analyzing high-throughput sequence data.
- To enable accurate identification and abundance estimation of non-host sequences.
- To provide a computationally efficient tool for pathogen detection in large datasets.
Main Methods:
- Development of a highly scalable parallel program named READSCAN.
- Utilizing a Beowulf compute cluster with 16 nodes for parallel processing.
- Testing READSCAN on a simulated dataset of 20.1 million reads.
Main Results:
- READSCAN accurately classified human and viral sequences.
- The program achieved classification in under 27 minutes on the simulated dataset.
- Demonstrated high scalability and efficiency on a modest compute cluster.
Conclusions:
- READSCAN is a powerful and efficient tool for identifying and quantifying non-host sequences in large-scale sequencing projects.
- The program's scalability and speed make it suitable for rapid analysis of potential pathogen-derived data.
- READSCAN contributes to advancing bioinformatics capabilities for genomic analysis and disease detection.

