Related Experiment Video
Updated: Sep 29, 2026

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
Kamino: fast proteome-wide variant calling for amino acid phylogenomics
Romain Derelle1, John A Lees2, Leonid Chindelevitch1
1MRC Centre for Global Infectious Disease Analysis, Department of Infectious Disease Epidemiology, School of Public Health, Imperial College London, London W12 , UK.
Abstract:
Amino acid-based phylogenetics usually relies on first clustering and aligning orthologous proteins. This approach is powerful but computationally demanding. Here, we present kamino, a reference- and alignment-free method that rapidly builds amino acid phylogenomic alignments directly from proteomes. As with similar algorithms, homologous regions are identified through shared sequences flanking variable regions. The method uses local changes in recoded k-mer occupancy to efficiently identify variable positions within these homologous regions and extract the corresponding pseudo-aligned sequences. It generates phylogenetically informative alignments across diverse prokaryotic and eukaryotic datasets. Phylogenetic analyses show that it accurately recovers Mycobacterium tuberculosis lineages, most curated GTDB taxa, and relationships consistent with published Drosophila and mammalian phylogenies, while producing signals broadly similar to BUSCO-based approaches. Runtimes are comparable to genome-based alignment-free methods and several orders of magnitude faster than classical marker-based pipelines, with moderate memory requirements. The method performs well across a broad range of divergence levels, from within-species comparisons to family-level prokaryotic and phylum-level eukaryotic datasets. kamino therefore provides a fast and simple route from proteomes to phylogenomic alignments across a broad range of evolutionary scales. The program is implemented in Rust and freely available at https://github.com/rderelle/kamino.
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

