Related Experiment Video
Updated: Aug 8, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Refined repetitive sequence searches utilizing a fast hash function and cross species information retrievals
1Department of Computer Science, University of Missouri, Columbia, USA. jeff@diglib1.cecs.missouri.edu
We developed a fast and memory-efficient algorithm to find small DNA sequences across multiple genomes. This tool helps researchers discover biological insights by analyzing repetitive DNA and associated annotations.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Small repetitive DNA sequences are crucial in various biological processes, from gene regulation in yeast to disease mechanisms in humans and virulence in bacteria.
- Current search algorithms for these sequences often yield numerous irrelevant results, hindering research efficiency.
Purpose of the Study:
- To develop a novel, time-efficient algorithm for locating small DNA sequences within multiple genomes.
- To integrate information retrieval algorithms and the Gene Ontology (GO) database for enhanced analysis of repetitive DNA sequences.
- To improve the accuracy and relevance of search results for genomic sequence analysis.
Main Methods:
- A hash function and search algorithm were developed to efficiently identify small DNA sequences across diverse genomes.
- Information retrieval techniques were applied to analyze cross-species conservation of repeat sequences.
- The Gene Ontology (GO) database was incorporated to refine search results with annotation data.
Main Results:
- The system demonstrates high performance, with an average search time of 1.147 seconds for an 8-base sequence across 3.224 GBases on 49 chromosomes.
- Searching with annotation terms significantly refines results, as shown in a yeast Pho4p binding site search.
- The algorithm efficiently handles multiple genomes without demanding significant main memory, outperforming other systems.
Conclusions:
- A time-efficient algorithm for locating small DNA segments and their associated annotation data has been developed.
- The algorithm effectively refines genome-wide search results by incorporating annotation data, reducing irrelevant hits.
- The developed algorithms are space-efficient, requiring minimal main memory, and are available upon request.
Related Concept Videos
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
Evolutionary Relationships through Genome Comparisons
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Sanger Sequencing
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...

