Related Experiment Video
Updated: Jun 26, 2026

16:17
The ITS2 Database
Published on: March 12, 2012
Sequence database search using jumping alignments
R Spang1, M Rehmsmeier, J Stoye
1German Cancer Research Center (DKFZ), Theoretical Bioinformatics, Heidelberg, Germany.
Summary
We developed a novel jumping alignment algorithm for protein sequence classification and remote homology detection. This method outperforms hidden Markov models in identifying distant protein relatives by utilizing both vertical and horizontal alignment information.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Sequence Analysis
Background:
- Established methods like profiles and hidden Markov models primarily focus on vertical information in multiple sequence alignments.
- There is a need for methods that can effectively exploit both vertical and horizontal information for robust protein classification.
- Detecting remote homologues is crucial for understanding protein function and evolution.
Purpose of the Study:
- To introduce a new algorithm for amino acid sequence classification and remote homology detection.
- To evaluate the performance of this new algorithm against established methods, specifically hidden Markov models.
Main Methods:
- Developed a 'jumping alignment' algorithm, extending the Smith-Waterman algorithm to compute local alignments between a single sequence and a multiple sequence alignment.
- The algorithm aligns a candidate sequence to a reference sequence within the multiple alignment, allowing the reference sequence to change with a penalty for jumps.
- Compared the discriminative quality of the jumping alignment algorithm against hidden Markov models using a subset of the SCOP database, assessing performance via false positive counts.
Main Results:
- The jumping alignment algorithm exploits both vertical and horizontal information in multiple sequence alignments, unlike methods focusing solely on vertical data.
- For moderate false positive counts (above five), the new algorithm demonstrated a considerably higher success rate in identifying true positives compared to hidden Markov models.
- The method effectively selects proteins belonging to a specific superfamily from a candidate database.
Conclusions:
- The novel jumping alignment algorithm offers a more balanced exploitation of multiple sequence alignment information.
- This approach shows superior performance in detecting remote protein homologues compared to hidden Markov models, particularly in challenging classification tasks.
- The algorithm represents a significant advancement in bioinformatics tools for protein family classification and evolutionary analysis.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Gene Duplication and Divergence
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Maxam-Gilbert Sequencing
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...

