fMLC: fast multi-level clustering and visualization of large molecular datasets
D Vu1, S Georgievska2, S Szoke1
1Bioinformatics group, Westerdijk Fungal Biodiversity Institute, 3584CT Utrecht, The Netherlands.
Motivation:
Despite successful applications of data clustering and visualization techniques in molecular sequence identification, current technologies still do not scale to large biological datasets.
Results:
We address this problem by a new multi-threaded tool, fMLC, primarily developed to cluster DNA sequences, that is supplemented with an interactive web-based visualization component, DiVE. fMLC enabled to compare, cluster and visualize 350K ITS fungal sequences at the species level. It took less than two hours to compare and cluster the dataset, which is twelve times faster than the time reported previously.
Availability And Implementation:
https://github.com/FastMLC/fMLC (doi: 10.5281/zenodo.926820).
Contact:
d.vu@westerdijkinstitute.nl or v.robert@westerdijkinstitute.nl.
More Related Videos
05:12ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
11:14Rapid High-throughput Species Identification of Botanical Material Using Direct Analysis in Real Time High Resolution Mass Spectrometry
Published on: October 2, 2016
Related Concept Videos
Applications of Molecular Taxonomy
Modern Molecular Taxonomy
Molecular Models
Evolutionary Relationships through Genome Comparisons
Distribution of Molecular Speeds
