Related Experiment Video
Updated: Jun 3, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Coverage bias in small molecule machine learning
Fleming Kretschmer1, Jan Seipp2, Marcus Ludwig1,3
1Chair for Bioinformatics, Institute for Computer Science, Friedrich Schiller University Jena, Jena, Germany.
Machine learning models for small molecules often lack coverage of biomolecular structures. This study introduces a new method to assess dataset coverage, improving model performance by guiding future data creation.
Area of Science:
- Computational chemistry
- cheminformatics
- machine learning
Background:
- Small molecule machine learning predicts properties from structures for applications like toxicity and drug discovery.
- End-to-end models are trending, but often overlook the domain of applicability and data coverage bias.
Purpose of the Study:
- To investigate the coverage of biomolecular structure space in large-scale datasets used for machine learning.
- To develop methods for assessing dataset representativeness and guiding future dataset creation.
Main Methods:
- Proposed a novel distance measure based on the Maximum Common Edge Subgraph (MCES) problem to quantify chemical similarity.
- Developed an efficient computational approach combining Integer Linear Programming and heuristic bounds to solve the MCES problem.
Main Results:
- Found that many widely-used datasets exhibit non-uniform coverage of biomolecular structures.
- This lack of uniform coverage limits the predictive power of machine learning models trained on these datasets.
Conclusions:
- Dataset coverage is a critical, often overlooked, factor in small molecule machine learning.
- The proposed MCES-based distance and divergence assessment methods can guide the creation of more representative datasets, enhancing model performance.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Drug Discovery: Overview
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
The Equilibrium Binding Constant and Binding Strength

