Related Experiment Video
Updated: Feb 28, 2026

13:22
Kinase Inhibitor Screening In Self-assembled Human Protein Microarrays
Published on: October 23, 2019
8.3K
Improving virtual screening predictive accuracy of Human kallikrein 5 inhibitors using machine learning models
Xingang Fang1, Sikha Bagui1, Subhash Bagui2
1Department of Computer Science, University of West Florida, Pensacola, FL 32514, United States.
Computational Biology and Chemistry
|June 12, 2017
Summary
Machine learning models can mine PubChem high throughput screening (HTS) data for drug discovery. A logistic regression model combined with Signature descriptors effectively identifies active molecules for Human kallikrein 5 (hK 5) inhibition.
Area of Science:
- * Computational chemistry and cheminformatics.
- * Application of machine learning in drug discovery.
Background:
- * High throughput screening (HTS) data from PubChem offers potential for machine learning-based small molecule mining.
- * Quantitative structure-activity relationship (QSAR) model development requires careful descriptor selection.
- * Understanding the interplay between descriptor selection, machine learning model choice, and target biomolecule characteristics is crucial for systematic strategy development.
Purpose of the Study:
- * To investigate the relationship between descriptor selection, machine learning models, and target biomolecule characteristics.
- * To develop a systematic strategy for descriptor and model selection in QSAR.
- * To evaluate the performance of different machine learning models using Signature descriptors for Human kallikrein 5 (hK 5) inhibition.
Main Methods:
- * Utilized Signature descriptors to generate a dataset from Human kallikrein 5 (hK 5) inhibition assay data.
- * Compared logistic regression, support vector machine, random forest, and k-nearest neighbor classification models.
- * Evaluated model performance using cross-validation and testing on a large HTS dataset (>200K structures).
Main Results:
- * The logistic regression model achieved high accuracy (98%) and precision (90%) in cross-validation for hK 5 inhibition.
- * The logistic regression model effectively eliminated over 99.9% of inactive structures from a large HTS dataset.
- * The combination of Signature descriptors and logistic regression demonstrated excellent predictive performance.
Conclusions:
- * The Signature descriptor and logistic regression model combination offers a feasible strategy for descriptor/model selection.
- * This approach shows promise for similar targets in drug discovery efforts.
- * Machine learning applied to HTS data can significantly enhance the efficiency of identifying potential drug candidates.
