Related Experiment Videos
Scaffold-shift uncertainty calibration in molecular activity prediction: A multi-target benchmark of coverage,
1Department of Computer Science, University of Illinois Springfield, Springfield, IL, 62704, USA.
Abstract:
Uncertainty estimates are increasingly used to prioritize molecules, yet their reliability under chemical scaffold shift is poorly characterized. We benchmarked prediction intervals and downstream selection policies across six human protein targets using 20,327 target-specific ChEMBL 37 compound records (19,972 globally unique ChEMBL molecule identifiers) and 7731 target-specific Bemis-Murcko scaffold instances, two model classes, and 20 scaffold-disjoint train/calibration/test partitions per target. We compared raw random-forest ensemble intervals, standard split conformal intervals, similarity-normalized conformal intervals, and similarity-binned local conformal intervals. Raw nominal 90% ensemble intervals achieved only 9.5%-11.4% molecule-weighted coverage. Conformal methods generally restored average coverage toward 90%, but results depended on target, partition, model, and whether molecules or scaffolds received equal weight. Conditional coverage averaged 82.3% in the least training-similar quartile and 78.7% in the highest-activity quartile. Similarity-adaptive methods did not uniformly improve the coverage-width trade-off. Across 240 matched target-partition-model comparisons, lower-confidence-bound selection reduced mean observed activity by 0.107 pActivity units, reduced top-decile hit rate by 0.063, and increased false optimism by 0.012, while selecting 2.16 additional scaffolds on average. Marginal calibration, conditional reliability, and decision utility are therefore distinct properties. Conformal calibration can repair severe interval undercoverage, but it does not automatically provide reliable extrapolation or a superior molecular-ranking policy.