CLAST: CUDA implemented large-scale alignment search tool.
Masahiro Yano1, Hiroshi Mori2, Yutaka Akiyama3
1Department of Biological Information, Graduate School of Bioscience and Biotechnology, Tokyo Institute of Technology, 2-12-1 M6-3, Ookayama, Meguro-ku, Tokyo, 152-8550, Japan. masayano@bio.titech.ac.jp.
A new tool, CLAST (CUDA implemented large-scale alignment search tool), rapidly analyzes microbial metagenomics data. It efficiently detects weak sequence similarities in massive datasets, overcoming limitations of existing methods for next-generation sequencing.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Metagenomics relies on nucleotide sequence similarity searching against databases.
- Next-generation sequencing generates vast amounts of data, often from unsequenced microbes.
- Existing tools struggle with large datasets and identifying weak similarities in novel microbial genomes.
Purpose of the Study:
- To develop a rapid sequence similarity search tool for large-scale metagenomic analyses.
- To address the limitations of current tools in handling massive next-generation sequencing data.
- To enable accurate detection of weak sequence similarities in microbial communities.
Main Methods:
- Development of CLAST (CUDA implemented large-scale alignment search tool).
- Utilizes NVIDIA Fermi architecture graphics processing units for high-speed computation.
- Supports both global and local alignment, with global alignment as default.
Main Results:
- CLAST is significantly faster than BLAST (~80.8x) and BLAT (~9.6x).
- Achieves high accuracy in assigning reads to taxonomic and functional groups using distant nucleotide sequences.
- Requires minimal main memory (<2 GB) and no preprocessed sequence databases.
Conclusions:
- CLAST offers high speed and sensitivity comparable to established tools like Bowtie 2, BLAST, BLAT, and FR-HIT.
- Eliminates the need for extensive database preprocessing or specialized hardware.
- Represents a powerful and practical approach for analyzing massive next-generation sequencing data in metagenomics.
More Related Videos
05:04Author Spotlight: Introduction to Active Probe Atomic Force Microscopy with Quattro-Parallel Cantilever Arrays
Published on: June 13, 2023
11:16Worm-align and Worm_CP, Two Open-Source Pipelines for Straightening and Quantification of Fluorescence Image Data Obtained from Caenorhabditis elegans
Published on: May 28, 2020
