Related Experiment Videos
Better prediction of sub-cellular localization by combining evolutionary and structural information.
1CUBIC, Department of Biochemistry and Molecular Biophysics, Columbia University, New York, New York 10032 , USA. nair@cubic.bioc.columbia.edu
Proteins
|November 25, 2003
Summary
Predicting protein localization is key to understanding function. This study develops a novel computational system using evolutionary and structural data, improving accuracy for extracellular and nuclear proteins.
Area of Science:
- Computational biology
- Protein bioinformatics
- Structural biology
Background:
- Protein sub-cellular localization is crucial for function.
- Existing prediction methods rely on amino acid composition or known targeting motifs.
- Limitations exist in current methods, especially for proteins lacking known motifs or homology.
Purpose of the Study:
- To develop an accurate in silico method for predicting protein sub-cellular localization.
- To explore the utility of evolutionary information and protein structure for localization prediction.
- To address the performance gap observed with known structure databases.
Main Methods:
- Utilized evolutionary information from multiple sequence alignments.
- Incorporated aspects of protein structure into prediction models.
- Developed a hybrid system combining statistical rules and neural networks.
- Created separate prediction systems for proteins with known versus unknown structures.
Main Results:
- Achieved over 65% four-state accuracy, outperforming composition-based methods.
- Demonstrated highest accuracy for extracellular and nuclear proteins.
- Observed significant performance degradation for methods trained on SWISS-PROT when applied to PDB data.
- Successfully annotated eukaryotic proteins in the Protein Data Bank (PDB) using the developed system.
Conclusions:
- A novel computational approach integrating evolutionary and structural data enhances protein localization prediction accuracy.
- Distinct prediction models are necessary for proteins with known versus unknown structures.
- The developed system holds promise for target selection in structural genomics and proteomic annotation.