Related Experiment Video
Updated: Jan 8, 2026

Electronic Tongue Generating Continuous Recognition Patterns for Protein Analysis
Published on: September 16, 2014
Self-supervised domain adaptation of protein language model based solely on positive enzyme-reaction pairs
Tomoya Okuno1, Naoaki Ono2, Md Altaf-Ul-Amin1
1Graduate School of Science and Technology, Nara Institute of Science and Technology, Ikoma, 630-0192, Nara, Japan.
Abstract:
There is growing interest in developing predictive models of enzyme catalytic properties that leverage activity data spanning diverse enzyme families. A fundamental challenge lies in the inherent biases of public biochemical databases. These databases predominantly catalog valid enzyme activities, rarely include negative instances, and report quantitative catalytic parameters for only a relatively small subset of enzymes. Such limitations pose a major obstacle to supervised learning of enzyme catalytic properties. One existing approach for model training involves generating synthetic negative enzyme-activity pairs by recombining existing enzymes and their activity information, particularly substrates or chemical reactions, that were not originally associated within datasets. However, it remains unclear whether the generated negative examples are truly inactive or merely unobserved active instances. To build a model that captures functional properties across diverse enzyme families while avoiding reliance on negative examples, this paper introduces a self-supervised domain adaptation methodology for pre-trained protein language models, solely based on positive enzyme-reaction pairs. The enzyme representations obtained from the adapted protein language model achieved superior or at least competitive performance compared to those from an existing method that relies on synthetic negatives, in both the turnover number prediction task for natural reactions of wild-type enzymes and the activity prediction task for family-wide enzyme-substrate specificity screening datasets. Overall, our approach represents a methodological advancement that eliminates the need for synthetic negatives and provides a scalable framework for leveraging the growing enzyme activity data in biochemical databases.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conservation of Protein Domains
Induced-fit Model
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding and Linkage

