Related Experiment Video
Updated: Jul 9, 2025

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
aaHash: recursive amino acid sequence hashing
Johnathan Wong1, Parham Kazemi1, Lauren Coombe1
1Canada's Michael Smith Genome Sciences Centre, BC Cancer, Vancouver, BC V5Z 4S6, Canada.
Motivation:
K-mer hashing is a common operation in many foundational bioinformatics problems. However, generic string hashing algorithms are not optimized for this application. Strings in bioinformatics use specific alphabets, a trait leveraged for nucleic acid sequences in earlier work. We note that amino acid sequences, with complexities and context that cannot be captured by generic hashing algorithms, can also benefit from a domain-specific hashing algorithm. Such a hashing algorithm can accelerate and improve the sensitivity of bioinformatics applications developed for protein sequences.
Results:
Here, we present aaHash, a recursive hashing algorithm tailored for amino acid sequences. This algorithm utilizes multiple hash levels to represent biochemical similarities between amino acids. aaHash performs ∼10× faster than generic string hashing algorithms in hashing adjacent k-mers.
Availability And Implementation:
aaHash is available online at https://github.com/bcgsc/btllib and is free for academic use.
Related Concept Videos
Amino acids
Leaky Scanning
Acid Halides to Amides: Aminolysis
In the first step of the aminolysis mechanism, the amine attacks the carbonyl carbon of the acyl chloride to form a tetrahedral intermediate. In the second step, the carbonyl group is re-formed with the elimination of a chloride...
Protein Organization
tRNA Activation
Mass Spectrometry of Amines
Amines undergo fragmentation through α cleavage, producing nitrogen-containing cations—iminium ions—and alkyl radicals. Mass spectra of aromatic and cyclic aliphatic amines exhibit...

