Related Experiment Videos
Predictive QSAR modeling based on diversity sampling of experimental datasets for the training and test set selection
Alexander Golbraikh1, Alexander Tropsha
1The Laboratory for Molecular Modeling, School of Pharmacy, University of North Carolina, Chapel Hill, NC 27599-7360, USA.
Journal of Computer-Aided Molecular Design
|December 20, 2002
Summary
Rational division of datasets improves Quantitative Structure Activity Relationship (QSAR) model predictive power. Employing diversity principles for training and test set selection enhances model accuracy for unseen compounds.
Area of Science:
- Cheminformatics
- Computational Chemistry
- Drug Discovery
Background:
- Predictive power is crucial for Quantitative Structure Activity Relationship (QSAR) models.
- QSAR models predict the properties of new compounds not included in the training data.
Purpose of the Study:
- To enhance QSAR model predictive power through rational dataset division.
- To establish criteria for selecting training and test sets based on molecular descriptor space.
Main Methods:
- Utilized molecular dataset diversity indices to quantify selection criteria.
- Employed sphere-exclusion algorithms for rational training and test set division.
- Validated the approach using multiple experimental datasets.
Main Results:
- QSAR models developed with the proposed method showed statistically superior predictive power.
- The rational selection approach outperformed random and activity ranking methods.
- Demonstrated improved accuracy in predicting biological activity for novel compounds.
Conclusions:
- Rational selection of training and test sets based on diversity principles is essential for robust QSAR modeling.
- This approach significantly improves the predictive accuracy of QSAR models.
- Recommends routine adoption of diversity-based selection in QSAR research.