Related Experiment Videos
A seqlet-based maximum entropy Markov approach for protein secondary structure prediction.
Qiwen Dong1, Xiaolong Wang, Lei Lin
1School of Computer Science and Technology, Harbin Institute of Technology, China. qwdong@insun.hit.edu.cn
Science in China. Series C, Life Sciences
|October 27, 2005
Summary
A new method uses protein secondary structure "seqlets" analogous to words to predict protein structures. This approach improves accuracy over existing methods and is available via a web server.
Area of Science:
- Computational biology
- Bioinformatics
- Structural biology
Background:
- Accurate protein secondary structure prediction is crucial for understanding protein function.
- Existing methods face challenges in capturing complex sequence-structure relationships.
Purpose of the Study:
- To develop a novel, accurate method for protein secondary structure prediction from amino acid sequences.
- To establish an organism-specific dictionary of protein secondary structure elements ('seqlets').
Main Methods:
- Protein secondary structure prediction formulated as a natural language processing problem (word segmentation and part-of-speech tagging).
- Extraction of 'seqlets' representing relationships between amino acid sequences and secondary structures.
- Utilizing a maximum entropy model and Viterbi algorithm for optimal prediction.
Main Results:
- The novel method outperforms the PHD method, achieving higher Q3 (3.9%) and SOV (4.6%) accuracy.
- Integration with BLAST local similarity search further enhances prediction accuracy.
- Achieved 78.9% Q3 and 77.1% SOV accuracy on CASP5 target proteins.
Conclusions:
- The seqlet-based approach offers a significant advancement in protein secondary structure prediction.
- The developed method provides a powerful tool for bioinformatics research.
- A web server is available for public use, facilitating broader application.