Related Experiment Video
Updated: Oct 21, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Computational Prediction of Compound-Protein Interactions for Orphan Targets Using CGBVS
Chisato Kanai1, Enzo Kawasaki1, Ryuta Murakami1
1Data Science Division, INTAGE Healthcare Inc., 2F NREG Midosuji Bldg., 3-5-7 Kawara-Machi, Chuo-ku, Osaka 541-0048, Japan.
Chemical genomics-based virtual screening (CGBVS) effectively predicts ligands for orphan G protein-coupled receptors (GPCRs). Multiple Sequence Alignment (MSA) protein descriptors enhance prediction accuracy, especially with ample related ligand data.
Area of Science:
- Computational chemistry
- Bioinformatics
- Drug discovery
Background:
- Artificial Intelligence (AI) and Machine Learning (ML) are increasingly used for in silico prediction of Compound-Protein Interactions (CPI).
- Chemical genomics-based virtual screening (CGBVS) is an AI technique utilizing support vector machine (SVM) for CPI prediction, known for accuracy and ease of use.
Purpose of the Study:
- To evaluate the efficacy of CGBVS in identifying ligands for orphan G protein-coupled receptors (GPCRs) lacking prior ligand information.
- To develop a method for assessing the applicability of CGBVS for predicting GPCR ligands.
Main Methods:
- Employed CGBVS with pairwise kernel-based SVM for predicting Compound-Protein Interactions (CPI).
- Utilized G protein-coupled receptor (GPCR) family data, specifically testing prediction on a virtual orphan GPCR with omitted ligand information during training.
- Assessed prediction accuracy using an applicability index and compared performance based on different protein and compound descriptors, including Multiple Sequence Alignment (MSA).
Main Results:
- Prediction accuracy varied significantly across different GPCRs.
- Models incorporating Multiple Sequence Alignment (MSA) as protein descriptors demonstrated superior overall prediction accuracy.
- The choice of protein descriptors had a greater impact on accuracy than the choice of compound descriptors.
- Prediction accuracy was positively correlated with the amount of available ligand information for related GPCRs.
Conclusions:
- CGBVS is a viable method for predicting ligands of orphan GPCRs.
- Multiple Sequence Alignment (MSA) is a key feature for improving prediction accuracy in CGBVS for GPCRs.
- The availability of related ligand data significantly influences the success of CGBVS in predicting GPCR ligands.
Related Concept Videos
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-Protein Interfaces

