Related Experiment Video
Updated: Feb 4, 2026

In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
Toward More Trustworthy QSAR: A Systematic Discussion on Data Set Partitioning
1School of Environmental Science and Engineering, Tianjin University, Tianjin 300072, China.
Data splitting methods and random seeds significantly impact QSAR model generalizability. Careful selection of test sets is crucial for reliable model evaluation, avoiding inflated performance metrics.
Area of Science:
- Quantitative Structure-Activity Relationship (QSAR) modeling
- cheminformatics
- computational chemistry
Background:
- QSAR model development is rapidly increasing.
- Concerns exist regarding the rigor of QSAR model evaluation, especially concerning data splitting strategies.
- The influence of data partitioning on model generalizability needs systematic investigation.
Purpose of the Study:
- To systematically assess the impact of random splits (RS), similarity-based splits (SS), and random seed variability on QSAR model generalizability.
- To evaluate these effects under scenarios with limited data (chemical screening) and ample data (standard modeling).
- To challenge the assumption that rational splits optimizing internal performance universally enhance model performance.
Main Methods:
- Utilized five datasets of varying sizes.
- Compared random splits (RS) and similarity-based splits (SS).
- Investigated the effect of random seed variability on model performance.
Main Results:
- Both data partitioning method and random seed selection substantially affect internal test performance, potentially misrepresenting true predictive ability.
- Similarity-based splits (SS) may improve internal performance but not always external generalizability.
- Under low sampling ratios, SS can underperform random splits (RS) on both internal and external tests.
- High variability in R-squared was observed across random seeds for internal tests on the smallest dataset, contrasting with lower variability on a fixed external dataset.
- Applicability domain (AD) filtering did not consistently reduce variability on the external dataset.
Conclusions:
- Test-set construction must align with real-world application scenarios for QSAR models.
- Researchers should avoid single or cherry-picked random seeds and unsuitable rational partitioning methods.
- Transparent, application-aligned partitioning protocols and AD methods are essential to prioritize true external generalizability over potentially inflated internal metrics.
Related Concept Videos
Design Example: Setting a Curve Using Design Data
Random and Systematic Errors
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Systematic Sampling Method
Systematic sampling is one of the simplest methods...
Propagation of Uncertainty from Systematic Error
Setting Time of Cement

