Related Experiment Video
Updated: Jun 17, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
The Fragility of Bioactivity Prediction: Rigorous Dataset Splits Expose the Illusion of ML Accuracy
Kisung Lee1, Galymzhan Moldagulov1,2, Bartosz A Grzybowski1,2
1Center for Algorithmic and Robotized Synthesis (CARS), Institute For Basic Science (IBS), Ulsan, Republic of Korea.
Abstract:
Machine learning (ML) has become the dominant approach for predicting biological activity from molecular structure, yet its true ability to generalize beyond known chemical space remains uncertain. Although numerous models have been proposed, recent work suggests that simple similarity-based methods can perform on par with far more sophisticated architectures, raising questions about whether current approaches genuinely learn transferable structure-activity relationships. Here, we systematically examine how different dataset-splitting strategies affect the performance of k-nearest neighbors (k-NN) and a representative set of modern ML models. Across all splitting regimes, k-NN performs comparably to state-of-the-art methods. More importantly, as dataset splits become increasingly stringent and test molecules move further out-of-distribution (OOD), the predictive accuracy of all models deteriorates sharply, approaching random guessing. This collapse persists even when standard fingerprints are augmented with 3D geometric descriptors, indicating that many previously reported accuracies-typically obtained under nonrigorous splitting strategies-are likely inflated. These results underscore the need for more rigorous evaluation standards and the development of representations capable of supporting true OOD generalization.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Improving Translational Accuracy
Improving Translational Accuracy
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...