Related Experiment Videos
Rapid protein domain assignment from amino acid sequence using predicted secondary structure.
Russell L Marsden1, Liam J McGuffin, David T Jones
1Bioinformatics Unit, Department of Computer Science, University College London, UK.
Protein Science : a Publication of the Protein Society
|November 21, 2002
Summary
Identifying protein domains without structural data is crucial. A new method, DomSSEA, shows promise by aligning predicted and observed secondary structures for accurate domain delineation.
Area of Science:
- Structural biology
- Bioinformatics
- Computational biology
Background:
- Determining protein domain content without structural information or sequence homology is a significant challenge.
- Accurate domain identification is vital for understanding protein function and evolution.
Purpose of the Study:
- To evaluate the effectiveness of domain delineation methods in the absence of sequence homology.
- To introduce and assess a novel method, DomSSEA, for continuous domain identification.
Main Methods:
- Comparison of baseline methods, Domain Guess by Size, and the newly developed DomSSEA algorithm.
- DomSSEA utilizes the alignment of predicted secondary structures against observed secondary structures from the CATH domain database.
- Sensitivity was measured by the number of correctly assigned top-scoring predictions.
Main Results:
- DomSSEA achieved a 73.3% success rate in correctly assigning domain numbers to a representative chain set.
- For multidomain proteins, DomSSEA correctly predicted domain number and boundary locations within a +/-20 residue margin for 24% of cases.
- Performance was contextualized against other assessed prediction methods.
Conclusions:
- The developed DomSSEA method offers a viable approach for automatic domain assignment, particularly when sequence homology is limited.
- Alignment of predicted and observed secondary structures provides a valuable basis for identifying continuous protein domains.
- Further refinement may improve accuracy for complex multidomain proteins.