Related Experiment Video
Updated: Jul 9, 2025

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
3.4K
aaHash: recursive amino acid sequence hashing
Johnathan Wong1, Parham Kazemi1, Lauren Coombe1
1Canada's Michael Smith Genome Sciences Centre, BC Cancer, Vancouver, BC V5Z 4S6, Canada.
Bioinformatics Advances
|November 29, 2023
Summary
A new hashing algorithm, aaHash, accelerates bioinformatics analyses of protein sequences by over 10x. This domain-specific approach accounts for biochemical similarities, improving speed and sensitivity for k-mer hashing.
Area of Science:
- Bioinformatics
- Computational Biology
- Proteomics
Background:
- K-mer hashing is crucial for bioinformatics but generic algorithms are inefficient for biological sequences.
- Existing methods do not leverage the specific alphabet and biochemical properties of amino acid sequences.
- There is a need for domain-specific hashing algorithms to enhance protein sequence analysis.
Purpose of the Study:
- To develop a novel hashing algorithm optimized for amino acid sequences.
- To improve the speed and sensitivity of bioinformatics applications for protein data.
- To address the limitations of generic string hashing in the context of protein sequences.
Main Methods:
- Introduction of aaHash, a recursive hashing algorithm specifically designed for amino acid sequences.
- aaHash employs multiple hash levels to capture biochemical similarities between amino acids.
- The algorithm is implemented and evaluated for its performance in k-mer hashing.
Main Results:
- aaHash demonstrates a significant speed improvement, performing approximately 10 times faster than generic string hashing algorithms for adjacent k-mers.
- The algorithm effectively utilizes the biochemical properties of amino acids for more efficient hashing.
- This leads to accelerated and potentially more sensitive analyses of protein sequences.
Conclusions:
- aaHash offers a substantial performance enhancement for k-mer hashing in protein sequence analysis.
- The domain-specific approach of aaHash improves upon generic hashing methods by considering amino acid biochemical properties.
- This tool has the potential to accelerate and refine various bioinformatics applications dealing with protein sequences.
Related Concept Videos
Amino acids
89.0K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
89.0K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Acid Halides to Amides: Aminolysis
2.8K
Aminolysis is a nucleophilic acyl substitution reaction, where ammonia or amines act as nucleophiles to give the substitution product. Acid halides react with ammonia, primary amines, and secondary amines to yield primary, secondary, and tertiary amides, respectively.
In the first step of the aminolysis mechanism, the amine attacks the carbonyl carbon of the acyl chloride to form a tetrahedral intermediate. In the second step, the carbonyl group is re-formed with the elimination of a chloride...
In the first step of the aminolysis mechanism, the amine attacks the carbonyl carbon of the acyl chloride to form a tetrahedral intermediate. In the second step, the carbonyl group is re-formed with the elimination of a chloride...
2.8K
Protein Organization
138.0K
Overview
138.0K
tRNA Activation
19.3K
Aminoacyl-tRNA synthetases are present in both eukaryotes and bacteria. Though eukaryotes have 20 different aminoacyl-tRNA synthetases to couple to 20 amino acids, many bacteria do not have genes for all of these aminoacyl-tRNA synthetases. Despite this, they still use all 20 amino acids to synthesize their proteins. For instance, some bacteria do not have the gene encoding the enzyme that couples glutamine with its partner tRNA. In these organisms, one enzyme adds glutamic acid to all of the...
19.3K
Mass Spectrometry of Amines
4.2K
In mass spectroscopy, amines undergo fragmentation to give parent ions with odd molecule weights. This observed mass spectrum follows the nitrogen rule: a molecule with an odd number of nitrogen atoms produces a parent ion with an odd molecular weight. The remaining fragments have an even mass.
Amines undergo fragmentation through α cleavage, producing nitrogen-containing cations—iminium ions—and alkyl radicals. Mass spectra of aromatic and cyclic aliphatic amines exhibit...
Amines undergo fragmentation through α cleavage, producing nitrogen-containing cations—iminium ions—and alkyl radicals. Mass spectra of aromatic and cyclic aliphatic amines exhibit...
4.2K

