Related Experiment Video
Updated: Oct 4, 2025

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency
Published on: June 16, 2011
The Statistics of k-mers from a Sequence Undergoing a Simple Mutation Process Without Spurious Matches
Antonio Blanca1, Robert S Harris2, David Koslicki1,2,3
1Department of Computer Science and Engineering, and The Pennsylvania State University, University Park, Pennsylvania, USA.
This study analyzes the statistical properties of k-mer mutations in biological sequences. We developed methods to estimate mutation rates (r) and assess sequence similarity using k-mer counts and Jaccard similarity.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- K-mer based methods are crucial in bioinformatics but lack thorough statistical understanding.
- Understanding the impact of sequence mutations on k-mer properties is essential for accurate data analysis.
Purpose of the Study:
- To statistically model the effects of independent nucleotide mutations on k-mer content.
- To derive methods for estimating mutation rates and sequence similarity.
- To provide practical applications in genome analysis and read alignment.
Main Methods:
- Developed a mathematical model for sequence mutation and its effect on k-mers.
- Derived expectation and variance for mutated k-mers, islands, and oceans.
- Formulated hypothesis tests and confidence intervals for mutation rates (r) using k-mer counts and Jaccard similarity.
Main Results:
- Provided theoretical framework for understanding k-mer changes under mutation.
- Derived novel statistical tests and confidence intervals for mutation rate estimation.
- Demonstrated practical utility in applications like Mash distance, read filtering, and alignment quality assessment.
Conclusions:
- The derived statistical methods enhance the reliability of k-mer based analyses in bioinformatics.
- This work offers a robust framework for quantifying sequence divergence and mutation.
- The applications highlight the practical impact of these statistical insights for genomic data interpretation.
More Related Videos
04:52Following the Dynamics of Structural Variants in Experimentally Evolved Populations
Published on: February 3, 2023
09:04Studying Ribonucleotide Incorporation: Strand-specific Detection of Ribonucleotides in the Yeast Genome and Measuring Ribonucleotide-induced Mutagenesis
Published on: July 26, 2018
Related Concept Videos
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Mutations in Microorganisms
Mutation, Gene Flow, and Genetic Drift
Mutations
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Spontaneous and Induced Mutations
Point and Frameshift Mutations