Related Experiment Video
Updated: Aug 29, 2026

In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
Rule-based labels versus experimental endpoints: A systematic assessment of generalisation in QSAR models for ADMET
Sandhiya Prabhakar1, Poojhashri Jayagopal1, J Jino Blessy1
1Department of Bioinformatics, Sri Ramachandra Faculty of Engineering and Technology, Sri Ramachandra Institute of Higher Education and Research, Chennai, Tamil Nadu, 600116, India.
Abstract:
Computational ADMET assessment is now a cornerstone of early drug discovery, yet QSAR classifiers are sometimes trained on physicochemical rule-derived labels rather than measured biological data. This practice raises a fundamental question: do strong internal performance statistics actually reflect predictive accuracy against experimental endpoints? To address this directly, we built Random Forest (RF) and XGBoost (XGB) classifiers for gastrointestinal (GI) absorption and blood-brain barrier (BBB) permeation using 995 fruit phytochemicals from 29 species, annotated with binary labels derived from published physicochemical thresholds; the same descriptors used to generate these labels were also supplied as model inputs, a design feature discussed further below. Stratified 80:20 partitioning was applied, and SMOTE oversampling was applied after splitting and restricted to the training fold to prevent data leakage. Hyperparameters were optimised by GridSearchCV with stratified five-fold cross-validation. All four model configurations attained essentially ceiling-level internal discrimination: AUC-ROC = 1.000 and MCC ≥ 0.980. Applying these models to an independent ChEMBL experimental dataset - Caco-2 apparent permeability values (n = 117) for GI absorption and logBB measurements (n = 102) for BBB permeation - caused AUC values to fall to 0.606-0.684 and MCC to 0.206-0.409, representing generalisation gaps of 0.32-0.39 AUC units. Y-randomisation (20 permutations, ΔAUC ≈ 0.50) indicated that the models encode a non-random structure-property signal rather than spurious correlation. A three-arm feature-representation comparison (physicochemical descriptors only, Morgan fingerprints only, and the combined representation used elsewhere in this study) found comparably near-ceiling internal performance across all three arms, indicating that the internal performance ceiling reflects the deterministic structure of the rule-derived labels rather than direct access to the label-generating descriptors specifically. A 95th-percentile Euclidean-distance applicability domain analysis showed complete structural coverage of the external compounds, indicating that structural out-of-domain extrapolation is unlikely to explain the observed gaps. These gaps are most plausibly attributed to the mismatch between computationally derived training labels and experimentally measured biological transport, although assay heterogeneity, transporter-mediated mechanisms, and other sources of biological complexity discussed in Section 4 are likely additional contributors. Topological polar surface area consistently emerged as the dominant predictive descriptor for both endpoints. These findings, drawn from a single dataset and two endpoints, support the case that external experimental validation deserves explicit attention in computational ADMET studies that rely on rule-derived training labels.
