Related Experiment Video
Updated: Mar 18, 2026

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency
Published on: June 16, 2011
Descriptive Statistics of the Genome: Phylogenetic Classification of Viruses
11 Mathematical Sciences Center, Tsinghua University , Beijing, China .
Abstract:
The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate families with the pairwise sequence alignment tool for similar species. The submitter's judgment is then used to determine the most likely species classification. The aim of this article is to show that this process can be automated into a fast, accurate, one-step process using the proposed alignment-free method and properly implemented machine learning techniques. We present a new family of alignment-free vectorizations of the genome, the generalized vector, that maintains the speed of existing alignment-free methods while outperforming all available methods. This new alignment-free vectorization uses the frequency of genomic words (k-mers), as is done in the composition vector, and incorporates descriptive statistics of those k-mers' positional information, as inspired by the natural vector. We analyze five different characterizations of genome similarity using k-nearest neighbor classification and evaluate these on two collections of viruses totaling over 10,000 viruses. We show that our proposed method performs better than, or as well as, other methods at every level of the phylogenetic hierarchy. The data and R code is available upon request.
Related Concept Videos
Size and Structure of Viral Genomes
Introduction to Virus
Viruses with RNA Genomes
Viruses of Archaea
Evolutionary Relationships through Genome Comparisons
Viral Structure

