Related Experiment Video
Updated: Nov 24, 2025

Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Minimally overlapping words for sequence similarity search
Martin C Frith1,2,3, Laurent Noé4, Gregory Kucherov5,6
1Artificial Intelligence Research Center, AIST, Tokyo, Japan.
Motivation:
Analysis of genetic sequences is usually based on finding similar parts of sequences, e.g. DNA reads and/or genomes. For big data, this is typically done via 'seeds': simple similarities (e.g. exact matches) that can be found quickly. For huge data, sparse seeding is useful, where we only consider seeds at a subset of positions in a sequence.
Results:
Here, we study a simple sparse-seeding method: using seeds at positions of certain 'words' (e.g. ac, at, gc or gt). Sensitivity is maximized by using words with minimal overlaps. That is because, in a random sequence, minimally overlapping words are anti-clumped. We provide evidence that this is often superior to acclaimed 'minimizer' sparse-seeding methods. Our approach can be unified with design of inexact (spaced and subset) seeds, further boosting sensitivity. Thus, we present a promising approach to sequence similarity search, with open questions on how to optimize it.
Availability And Implementation:
Software to design and test minimally overlapping words is freely available at https://gitlab.com/mcfrith/noverlap.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Sequences

