Related Experiment Videos
Embedding strategies for effective use of information from multiple sequence alignments
1Howard Hughes Medical Institute, Basic Sciences Division, Fred Hutchinson Cancer Research Center, Seattle, Washington 98104, USA. henikoff@howard.fhcrc.org
Protein Science : a Publication of the Protein Society
|March 1, 1997
Summary
This study introduces a novel method for protein sequence analysis, enhancing database searches by embedding conserved region information. This strategy improves the detection of distant evolutionary relationships in protein families.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Detecting distant evolutionary relationships between protein sequences is crucial for understanding protein function and evolution.
- Traditional methods often struggle with identifying remote homologs due to sequence divergence.
Purpose of the Study:
- To develop and evaluate a new strategy for enhancing protein sequence database searches.
- To improve the detection of distant relationships by incorporating multiple sequence alignment information.
Main Methods:
- Enriching a representative protein sequence by replacing conserved regions with position-specific scoring matrices (PSSMs) or consensus residues.
- Testing PSSM-embedded queries with a specialized Smith-Waterman algorithm.
- Evaluating consensus residue-embedded queries with standard search tools like BLAST and FASTA.
Main Results:
- PSSM-embedded queries combined with a specialized Smith-Waterman algorithm yielded the best overall search performance.
- Embedding consensus residues improved performance with widely used single-sequence search programs (BLAST, FASTA).
- This approach effectively balances conserved motif information with single-sequence data.
Conclusions:
- Embedding PSSMs or consensus residues into representative sequences is an effective strategy for improving protein sequence similarity searches.
- The method enhances the ability to detect distant protein relationships by leveraging multiple alignment data.
- This technique offers improved performance for various sequence search algorithms.