Reduced alphabet motif methodology for GPCR annotation
Rajeev Gangal1, K Krishna Kumar
1Insilico Consulting, 402, Citi Centre, 39/2, Erandwane, Karve Road, Pune, Maharashtra, India.
Journal of Biomolecular Structure & Dynamics
|October 17, 2007
Summary
This study introduces a novel method for classifying G-protein coupled receptors (GPCRs) using reduced protein alphabets and regular expression motifs. This approach accurately identifies GPCR families, overcoming limitations of existing machine learning and sequence similarity methods.
Area of Science:
- Computational biology
- Bioinformatics
- Genomics
Background:
- G-protein coupled receptors (GPCRs) classification is computationally challenging.
- Existing methods like machine learning and sequence motif analysis have limitations such as class imbalance, lack of multi-class classification, and over-reliance on primary sequence similarity.
Purpose of the Study:
- To develop and evaluate a novel motif discovery and application methodology for GPCR classification.
- To overcome the limitations of existing machine learning and sequence similarity-based approaches.
- To accurately classify GPCRs, including orphan receptors, and predict their families and subfamilies.
Main Methods:
- Utilizing a reduced protein functional alphabet representation where functionally similar residues share symbols.
- Generating regular expression motifs via ClustalW based multiple sequence alignment with an identity matrix.
- Employing this method for sequences without necessarily sharing common ancestry, avoiding evolutionary matrices like BLOSUM and PAM.
Main Results:
- The reduced alphabet motifs accurately classify known GPCR proteins, achieving results comparable to PRINTS and PROSITE.
- Demonstrated high accuracy with no false negatives and few false positives for well-known GPCRs in SWISSPROT.
- Successfully classified most known GPCR classes, even with limited representative sequences, and predicted multiple classes for certain sequences.
- Annotated 695 orphan receptors, identifying 121 belonging to Family A.
Conclusions:
- The proposed methodology offers an effective and versatile approach for GPCR classification.
- It overcomes key limitations of traditional machine learning and sequence similarity methods.
- The developed web interface provides a practical tool for predicting GPCR families and subfamilies.


