Centroid based clustering of high throughput sequencing reads based on n-mer counts

Alexander Solovyov1, W Ian Lipkin

  • 1Center for Infection and Immunity, Columbia University, New York, NY, 10032, USA. avs2132@columbia.edu.

BMC Bioinformatics
|September 10, 2013
PubMed
Summary

Alignment-free sequence clustering using word counts is efficient for computational biology tasks. This method, particularly with soft expectation maximization, enhances clustering accuracy for high-throughput sequencing analysis.