Related Experiment Video
Updated: Jul 27, 2025

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Matchtigs: minimum plain text representation of k-mer sets
Sebastian Schmidt1, Shahbaz Khan2, Jarno N Alanko3,4
1Department of Computer Science, University of Helsinki, Helsinki, Finland. sebastian.schmidt@helsinki.fi.
Abstract:
We propose a polynomial algorithm computing a minimum plain-text representation of k-mer sets, as well as an efficient near-minimum greedy heuristic. When compressing read sets of large model organisms or bacterial pangenomes, with only a minor runtime increase, we shrink the representation by up to 59% over unitigs and 26% over previous work. Additionally, the number of strings is decreased by up to 97% over unitigs and 90% over previous work. Finally, a small representation has advantages in downstream applications, as it speeds up SSHash-Lite queries by up to 4.26× over unitigs and 2.10× over previous work.
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Karyotyping
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
¹H NMR Signal Multiplicity: Splitting Patterns
Sanger Sequencing
DNA Base Pairing

