Related Experiment Video
Updated: Oct 4, 2026

Defining Substrate Specificities for Lipase and Phospholipase Candidates
Published on: November 23, 2016
Information Leakage in Enzyme Substrate Prediction
Vahid Atabaigi Elmi1,2, Roman Joeres1,2, Olga V Kalinina1,2,3
1Drug Bioinformatics, Helmholtz Institute for Pharmaceutical Research Saarland, Campus E8.1, 66123 Saarbruecken, Saarland, Germany.
Motivation:
Enzymes are essential catalysts in many cellular processes. Understanding their interactions with small molecules, such as regulators, cofactors, and most importantly, substrates, is crucial for understanding the biochemical processes that occur in cells. Correctly interpreting the roles of small molecules that interact with enzymes is key to elucidating enzyme function. Recently, enzyme-small molecule interaction prediction has attracted growing interest from computational methods, especially deep learning. As a result, researchers have published several datasets and numerous models with remarkable performance.
Results:
In this work, we critically examine one of the most popular datasets and four models trained on it, identifying similarity-induced information leakage that may overinflate reported model performance in out-of-distribution applications. We show that the inspected models are susceptible to information leakage, and their performance is not better than statistical baselines when the leakage is removed. Furthermore, controlling for leakage due to protein sequence similarity alone still yields high predictive performance, whereas performance decreases substantially when common substrates shared between training and test sets are removed from one of them, and approaches random prediction when structurally similar ligands are separated between splits. Thus, the investigated models generalize considerably better across enzyme sequence space than across ligand chemical space, indicating that their reported performance depends strongly on the small-molecule similarity between training and evaluation data.
Availability And Implementation:
All data splits we calculated in this study are available on Zenodo (DOI: 10.5281/zenodo.18786394). The code is available on GitHub and backed up on Zenodo (DOI: 10.5281/zenodo.18788609).
Supplementary Information:
Supplementary information is available at Bioinformatics online.
Related Concept Videos
Enzymes
Enzyme deficiencies can often translate into life-threatening diseases. For example, a genetic abnormality resulting in the deficiency of the enzyme G6PD...
Induced-fit Model
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical characteristics of...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Turnover Number and Catalytic Efficiency
Chymotrypsin is a pancreatic enzyme that breaks down proteins during digestion. The...
Enzyme Inhibition
Catalytically Perfect Enzymes

