Related Experiment Video
Updated: May 28, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Sequence-based enzyme catalytic domain prediction using clustering and aggregated mutual information content.
1Division of Experimental Hematology and Cancer Biology, Cincinnati Children's Hospital Medical Center, 3333 Burnet Avenue, Cincinnati, Ohio 45229, USA. kwangmin.choi.phd@gmail.com
This study introduces a novel in silico method for identifying enzyme active sites using sequence clustering and information theory. The approach efficiently predicts catalytic residues, aiding in enzyme function prediction for large-scale genome projects.
Area of Science:
- Bioinformatics
- Computational Biology
- Enzymology
Background:
- Experimental methods for enzyme active site identification are costly and time-consuming.
- High-throughput in silico methods are needed for analyzing vast protein sequence data.
- Accurate prediction of catalytic residues is crucial for enzyme function prediction.
Purpose of the Study:
- To develop a novel, accurate, and computationally efficient in silico method for predicting enzyme active sites.
- To address the challenges of varying sequence similarity in enzyme datasets.
- To facilitate enzyme function prediction in large-scale genomic studies.
Main Methods:
- Enzyme sequences are clustered based on functional categories (EC labels).
- Sequence graphs are constructed, utilizing graph properties like biconnected components and articulation points to identify common sequence segments.
- An information-theoretic approach, the aggregated column related scoring scheme, is applied to aligned subsequences in common regions to identify potential active sites.
Main Results:
- The combined clustering and aggregated information content scoring method successfully identified known catalytic sites in Escherichia coli K12 enzymes.
- The method demonstrated accuracy in predicting potential active sites.
- The approach is computationally efficient, with graph properties computed in linear time relative to graph edges.
Conclusions:
- The proposed sequence-based method offers an accurate and efficient solution for identifying enzyme active sites.
- This computational approach can significantly aid in the analysis of enzyme sequences from numerous genome projects.
- The method provides a valuable tool for advancing enzyme function prediction and understanding.
More Related Videos
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Induced-fit Model
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical characteristics of...

