Related Experiment Video
Updated: Mar 21, 2026

09:26
DNA-Tethered RNA Polymerase for Programmable In vitro Transcription and Molecular Computation
Published on: December 29, 2021
4.9K
genCRC32: collision-free CRC32-based hashing of DNA sequences
Pavel Beran1, Michael Rost1, Kristina Beranová1
1Faculty of Agriculture and Technology, University of South Bohemia in České Budějovice, České Budějovice 370 05, Czech Republic.
Bioinformatics Advances
|March 20, 2026
Summary
genCRC32 offers collision-free hashing for DNA k-mers up to length 16, crucial for bioinformatics. This method ensures accurate sequence analysis without performance loss.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Efficient and collision-free hashing of DNA sequences is critical for bioinformatics tools like genome assembly and sequence alignment.
- Traditional hashing methods often lead to collisions, compromising the accuracy and performance of downstream analyses.
- The practical limits of 32-bit hashing are reached for DNA k-mers up to length 16, necessitating improved hashing strategies.
Purpose of the Study:
- To evaluate genCRC32 as a hashing primitive for DNA sequences.
- To assess its collision behavior, bucket balance, sensitivity to single-base changes, and speed.
- To determine its suitability for collision-free mapping of DNA k-mers up to length 16.
Main Methods:
- Developed genCRC32 by combining a gen32 preprocessing step with CRC32 hashing.
- Identified eight specific CRC32 polynomials that guarantee collision-free hashing for DNA k-mers up to length 16.
- Conducted extensive empirical evaluations and benchmark tests comparing genCRC32 with established hashing methods.
Main Results:
- genCRC32 achieved zero collisions for all DNA k-mers up to length 16, ensuring a one-to-one mapping.
- The preprocessing step introduced minimal computational overhead, maintaining hashing performance comparable to MurmurHash3 and xxHash32.
- The method demonstrated good bucket balance and sensitivity to single-base changes.
Conclusions:
- genCRC32 is an effective and efficient hashing method for DNA sequences, particularly for k-mers up to length 16.
- Its collision-free nature and comparable performance make it a valuable primitive for bioinformatics applications.
- The publicly available Go implementation facilitates integration into existing bioinformatics workflows.
Related Concept Videos
Sanger Sequencing
777.7K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
777.7K
Maxam-Gilbert Sequencing
13.6K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
13.6K
DNA as a Genetic Template
28.5K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
28.5K
Overview of DNA Repair
34.8K
In order to be passed through generations, genomic DNA must be undamaged and error-free. However, every day, DNA in a cell undergoes several thousand to a million damaging events by natural causes and external factors. Ionizing radiation such as UV rays, free radicals produced during cellular respiration, and hydrolytic damage from metabolic reactions can alter the structure of DNA. Damages caused include single-base alteration, base dimerization, chain breaks, and cross-linkage.
Chemically...
Chemically...
34.8K
Homologous Recombination
65.1K
The basic reaction of homologous recombination (HR) involves two chromatids that contain DNA sequences sharing a significant stretch of identity. One of these sequences uses a strand from another as a template to synthesize DNA in an enzyme-catalyzed reaction. The final product is a novel amalgamation of the two substrates. To ensure an accurate recombination of sequences, HR is restricted to the S and G2 phases of the cell cycle. At these stages, the DNA has been replicated already and the...
65.1K
DNA Isolation
46.1K
DNA isolation protocols can be fast and straightforward or complex and time-consuming depending on the type and quality of DNA required for further processing. For example, plasmid DNA extraction is a bit more complicated than genomic DNA extraction because of the need for an appropriate lysis method to separate plasmid DNA from gDNA during isolation. However, for specific applications, such as long-range DNA sequencing that require a good yield of high- quality DNA samples, we need to follow...
46.1K

