Related Experiment Video
Updated: Aug 29, 2026

A Pipeline to Investigate the Structures and Signaling Pathways of Sphingosine 1-Phosphate Receptors
Published on: June 8, 2022
Validation subset selections for extrapolation oriented QSPAR models
Csaba Szántai-Kis1, István Kövesdi, György Kéri
1Cooperative Research Center, Semmelweis University, Pf 131, Budapest 5, Hungary, 1367. szacsa@rezso.sote.hu
Abstract:
One of the most important features of QSPAR models is their predictive ability. The predictive ability of QSPAR models should be checked by external validation. In this work we examined three different types of external validation set selection methods for their usefulness in in-silico screening. The usefulness of the selection methods was studied in such a way that: 1) We generated thousands of QSPR models and stored them in 'model banks'. 2) We selected a final top model from the model banks based on three different validation set selection methods. 3) We predicted large data sets, which we called 'chemical universe sets', and calculated the corresponding SEPs. The models were generated from small fractions of the available water solubility data during a GA Variable Subset Selection procedure. The external validation sets were constructed by random selections, uniformly distributed selections or by perimeter-oriented selections. We found that the best performing models on the perimeter-oriented external validation sets usually gave the best validation results when the remaining part of the available data was overwhelmingly large, i.e., when the model had to make a lot of extrapolations. We also compared the top final models obtained from external validation set selection methods in three independent and different sizes of 'chemical universe sets'.
Related Concept Videos
Detection of Gross Error: The Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Distributions to Estimate Population Parameter
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...