Related Experiment Videos
Judging the significance of multiple linear regression models
David J Livingstone1, David W Salt
1ChemQuest, Sandown, UK. davel@chemquest.uk.com
Journal of Medicinal Chemistry
|February 4, 2005
Summary
Statistical tests for multiple linear regression (MLR) models are often inappropriate due to selection bias. New critical values (Fmax) derived from random number experiments provide a more accurate assessment of model significance in drug discovery.
Area of Science:
- Cheminformatics
- Computational Chemistry
- Drug Discovery
Background:
- Multiple linear regression (MLR) is widely used in cheminformatics to model biological activity.
- Calculating numerous molecular descriptors and applying variable selection is common practice.
- Standard statistical tests for MLR significance are often misapplied in this context.
Purpose of the Study:
- To address the issue of "selection bias" in MLR models built from large descriptor sets.
- To introduce a more appropriate method for assessing the statistical significance of these models.
- To improve the reliability of QSAR (Quantitative Structure-Activity Relationship) models.
Main Methods:
- Generating large numbers of molecular descriptors.
- Applying variable selection techniques.
- Constructing MLR models.
- Conducting regression experiments with random numbers to determine critical values (Fmax).
Main Results:
- Established that standard statistical tests are inappropriate for MLR models with selection bias.
- Determined critical Fmax values through simulations with random data.
- Demonstrated a method for more accurate significance assessment of MLR models.
Conclusions:
- The "selection bias" inherent in descriptor selection invalidates traditional significance tests for MLR models.
- The proposed Fmax values offer a statistically sound basis for evaluating the significance of QSAR models.
- This approach enhances the reliability of predictive models in drug discovery and development.