Related Experiment Video
Updated: Mar 24, 2026

Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Read-Consistent Minimum Unique Substrings: A Parameter-Free, Linear-Time Framework for Genomic Sequence
Abstract:
Fixed-length k-mers have been the standard unit of genomic sequence representation for over two decades. However, they impose a uniform resolution on genomes whose complexity varies across loci. We introduce Minimum Unique Substrings (MUSs), variable-length sequence units defined by the local uniqueness structure of the genome rather than predefined parameters. We first extend MUS theory from single contiguous strings to fragmented sequencing reads by formalizing a definition of uniqueness that is consistent with these reads. Next, we present a linear-time extraction algorithm that runs in O(n) time using the generalized suffix tree. In this context, we introduce outpost nodes, topological anchors within the suffix tree that accurately localize MUS boundaries in fragmented sequencing reads. Finally, we empirically characterize the distributions of MUS lengths in E. coli K-12 and human chromosome 11. Our results demonstrate that MUS lengths naturally mirror genomic architectural complexity without the need for user-defined parameters. Notably, the MUS framework achieves 100% unique positional coverage with a mean length of only 36.08 bp. In contrast, fixed-length k=61 coverage reaches only 69.4%, despite being 1.69 times the MUS average. We show that increasing k from 21 to 61 triples the unique k-mer count from 2.35M to 6.86M. This k-paradox occurs because repetitive sequences are fragmented into spuriously unique tokens without improving true genomic resolution. MUSs escape this artifact entirely by adapting dynamically to local sequence complexity. These results establish MUSs as a biologically grounded, computationally tractable foundation for parameter-free genome assembly, repeat characterization, and alignment-free genomics.
Related Concept Videos
Modern Molecular Taxonomy
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Sanger Sequencing
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Evolutionary Relationships through Genome Comparisons

