Auditing widely used biomolecular benchmarks reveals systematic data inconsistencies
Maximilian G Schuh1, Aleksandra Daniluk1, Stephan A Sieber1
1TUM School of Natural Sciences, Department of Biosciences, Chair of Bioorganic Chemistry, Center for Functional Protein Assemblies, Technical University of Munich (TUM) Ernst-Otto-Fischer-Str. 8 85748 Garching Germany stephan.sieber@tum.de.
Abstract:
Machine learning accelerates molecular discovery and relies heavily on standardised benchmark datasets to evaluate computational performance. To learn transferable structure-activity relationships (SARs), models must be trained on datasets that accurately reflect chemical and biological realities without introducing artificial redundancies. Although standardised benchmarks are ubiquitous, the integrity of their underlying data is often assumed rather than rigorously verified. Currently, there is a lack of systematic and quantitative assessments of hidden data leakage and structural inconsistencies across widely used biomolecular benchmarks. Here, we demonstrate that several prominent chemical and biochemical benchmarking suites have pervasive cross-split contamination, unresolved label conflicts, and severe structural redundancies. By auditing over fifty dataset configurations, we reveal that tasks previously considered robust evaluation environments often have hidden flaws, particularly in drug-target interaction datasets with extensive protein overlap. Controlled noise-injection experiments show that these artefacts can systematically bias benchmark metrics. A complementary counterfactual leaderboard analysis further shows that leading-model conclusions can change when the audited composition of the test set is altered, particularly when label conflicts are enriched. These findings suggest that many leaderboard and leading-model conclusions should be interpreted in light of benchmark composition and data quality, rather than as automatic evidence of robust methodological superiority. Our findings highlight the importance of auditing evaluations, reporting chemically meaningful test-set composition, and using leakage-resistant data splits to accurately measure model generalisability. Moving from uncritical leaderboard optimisation to rigorous dataset auditing will yield more reliable computational tools. Ultimately, this will ensure that artificial intelligence models can be reliably translated into real-world applications in therapeutic design.

