Related Experiment Video
Updated: Jan 13, 2026

06:38
Pattern-based Search of Epigenomic Data Using GeNemo
Published on: October 8, 2017
5.4K
A rapid access motif database (RAMdb) with a search algorithm for the retrieval patterns in nucleic acids or protein
1CIT12 (Centre Interuniversitaire de Traitement de l'Information), Universite Paris, France.
Summary
This study introduces a novel bioinformatics algorithm for rapid sequence searching in large biomolecule databases. The new method enhances speed and efficiency for pattern discovery in nucleic acid and protein sequences.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biomolecule database management requires efficient sequence retrieval methods.
- Existing homology search methods can be time-consuming for large datasets.
- Previous work involved preprocessed codification of databases for faster searching.
Purpose of the Study:
- To develop a new search algorithm for rapid sequence retrieval from biomolecule databases.
- To improve upon existing methods for homology search and pattern discovery.
- To create a system applicable to both nucleic acid and protein sequences.
Main Methods:
- A codification structure interfaced with biomolecule database management packages.
- A new search algorithm utilizing recognition of rarest strings and sorted list intersection.
- Application of the system to find patterns in databanks and large sequence sets.
Main Results:
- The new algorithm significantly speeds up the search for query patterns in large datasets.
- The system demonstrates applicability to both nucleic acid and protein sequence analysis.
- Comparison with existing methods confirms the enhanced search speed.
Conclusions:
- The presented codification structure and search algorithm offer a faster and more efficient approach to sequence retrieval.
- This method is valuable for pattern discovery and analysis in large-scale biological sequence data.
- The system provides a significant advancement for bioinformatics research and applications.
Related Concept Videos
Modern Molecular Taxonomy
576
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
576
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Proteomics
9.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
9.3K
DNA Microarrays
20.6K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
20.6K

