Auditing widely used biomolecular benchmarks reveals systematic data inconsistencies
Maximilian G Schuh1, Aleksandra Daniluk1, Stephan A Sieber1
1TUM School of Natural Sciences, Department of Biosciences, Chair of Bioorganic Chemistry, Center for Functional Protein Assemblies, Technical University of Munich (TUM) Ernst-Otto-Fischer-Str. 8 85748 Garching Germany stephan.sieber@tum.de.
Chemical Science
|August 1, 2026
Summary
Machine learning for drug discovery needs reliable data. This study found hidden flaws like data leakage and inconsistencies in common benchmarks, impacting model performance evaluations.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
- Bioinformatics
Background:
- Machine learning (ML) accelerates molecular discovery, relying on standardized benchmark datasets for performance evaluation.
- Accurate structure-activity relationship (SAR) learning requires datasets reflecting chemical/biological realities without artificial redundancies.
- Data integrity in benchmarks is often assumed, lacking systematic, quantitative assessments for hidden data leakage and structural inconsistencies.
Purpose of the Study:
- To systematically and quantitatively assess hidden data leakage, label conflicts, and structural redundancies in widely used biomolecular benchmark datasets.
- To reveal the impact of data artefacts on the evaluation of machine learning models for drug discovery.
- To emphasize the need for rigorous dataset auditing and improved data splitting strategies.
Main Methods:
- Auditing over fifty dataset configurations across prominent chemical and biochemical benchmarking suites.
- Controlled noise-injection experiments to assess the impact of artefacts on benchmark metrics.
- Counterfactual leaderboard analysis to evaluate the sensitivity of model performance to test set composition.
Main Results:
- Pervasive cross-split contamination, unresolved label conflicts, and severe structural redundancies were found in several prominent benchmark suites.
- Data artefacts were particularly prevalent in drug-target interaction datasets with significant protein overlap.
- Noise-injection and leaderboard analyses demonstrated that these artefacts systematically bias benchmark metrics and can alter leading-model conclusions.
Conclusions:
- Many current benchmark evaluations and leading-model conclusions may not accurately reflect true model generalisability due to data quality issues.
- Rigorous dataset auditing, transparent reporting of test set composition, and leakage-resistant data splits are crucial for reliable ML model evaluation.
- Transitioning from uncritical leaderboard optimization to dataset auditing will enhance the reliability of computational tools for therapeutic design.

