Related Experiment Video
Updated: May 13, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Time-split cross-validation as a method for estimating the goodness of prospective prediction
1Cheminformatics Department, Merck Research Laboratories, Rahway, New Jersey 07065, USA. sheridan@merck.com
Journal of Chemical Information and Modeling
|March 26, 2013
Summary
Time-split selection improves quantitative structure-activity relationship (QSAR) model validation by providing a more realistic estimate of predictive performance compared to random or leave-class-out methods.
Area of Science:
- Quantitative structure-activity relationship (QSAR) modeling
- Computational chemistry
- Drug discovery
Background:
- Cross-validation is essential for validating QSAR models, assessing their self-consistency and predictive ability.
- Traditional cross-validation methods like random selection can yield overly optimistic predictions.
- Other methods, such as leave-class-out, may be too pessimistic.
Purpose of the Study:
- To evaluate different cross-validation strategies for QSAR model validation.
- To determine which cross-validation method best reflects true prospective prediction accuracy.
- To propose an improved standard for QSAR model cross-validation.
Main Methods:
- Comparison of R-squared (R²) values obtained from different test set selection strategies.
- Implementation of random selection, leave-class-out analog, and time-split selection.
- Analysis of how test set composition influences the estimation of model predictivity.
Main Results:
- Time-split selection yields R² values that more closely align with true prospective prediction.
- Random selection consistently overestimates model predictivity.
- The analog of leave-class-out selection tends to underestimate model predictivity.
Conclusions:
- Time-split selection offers a more reliable assessment of QSAR model predictive performance.
- It is recommended to incorporate time-split selection alongside random selection in QSAR model building.
- This approach enhances the robustness and trustworthiness of QSAR model validation.
More Related Videos
Related Concept Videos
Prediction Intervals
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Survival Tree
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a survival tree begins...
Building a Survival Tree
Constructing a survival tree begins...
Sensitivity, Specificity, and Predicted Value
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
Goodness-of-Fit Test
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
End Point Prediction: Gran Plot
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting the...
For potentiometric titration, the Gran plot is created by plotting the...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
