Related Experiment Video
Updated: Mar 29, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Enzyme mechanism prediction: a template matching problem on InterPro signature subspaces
Hamse Y Mussa1, Luna De Ferrari2, John B O Mitchell3
1EaStCHEM School of Chemistry and Biomedical Sciences Research Complex, University of St Andrews, North Haugh, St Andrews, KY16 9ST, Scotland, UK. mussax021@gmail.com.
A simple k Nearest Neighbour (kNN) algorithm accurately predicts enzyme chemical mechanisms using InterPro sequence signatures. This study explains the high accuracy by showing distinct enzyme features in a large, sparse feature space, simplifying classification.
Area of Science:
- Biochemistry
- Bioinformatics
- Computational Biology
Background:
- Previous work demonstrated high accuracy in predicting enzyme chemical mechanisms using a k Nearest Neighbour (kNN) rule with InterPro sequence signatures.
- The kNN algorithm's sensitivity to training data errors and small datasets was a concern in prior studies.
- The current study re-analyzed the dataset and prediction results to understand the observed high classification performance.
Purpose of the Study:
- To explain the remarkable classification performance of a k Nearest Neighbour (kNN) rule for predicting enzyme chemical mechanisms.
- To investigate the underlying reasons for the high accuracy despite a large and sparse feature space.
Main Methods:
- Utilized a k Nearest Neighbour (kNN) rule with k=1 (k1NN).
- Employed 321 InterPro sequence signatures as enzyme features.
- Re-analyzed a dataset of 248 enzymes with 71 MACiE database mechanism labels.
Main Results:
- Enzymes with different chemical mechanisms were found in barely overlapping subspaces within the feature space.
- The selected InterPro features contained sufficient information for accurate enzymatic mechanism classification.
- The classification problem was effectively reduced to a simple look-up exercise due to feature distinctiveness.
Conclusions:
- The study explains the "anomaly" of a basic kNN algorithm achieving high performance in a large, sparse feature space for enzyme mechanism prediction.
- InterPro signatures were confirmed as critical for accurate enzyme mechanism prediction.
- Simple rules were suggested for inductively predicting predefined mechanisms in novel enzymes.
More Related Videos
Related Concept Videos
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
¹H NMR: Complex Splitting
Splitting diagrams or splitting tree diagrams are routinely used to depict such complex couplings. While drawing splitting diagrams, the splitting with the larger coupling constant is usually applied...
Protein-protein Interfaces
Protein Complexes with Interchangeable Parts
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
Interpreting ¹H NMR Signal Splitting: The (n + 1) Rule

