Related Experiment Video
Updated: Jun 19, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Partially-supervised protein subclass discovery with simultaneous annotation of functional residues
Benjamin Georgi1, Jörg Schultz, Alexander Schliep
1Max Planck Institute for Molecular Genetics, Dept, of Computational Molecular Biology, Ihnestrasse 73, 14195 Berlin, Germany. bgeorgi@mail.med.upenn.edu
This study introduces a novel partially-supervised learning method to identify protein domain subfamilies and predict functional residues. This approach enhances understanding of substrate specificity in protein families like phosphatases and WW domains.
Area of Science:
- Computational biology
- Bioinformatics
- Structural biology
Background:
- Identifying functional subfamilies within protein domain families is crucial for understanding substrate specificity.
- Clustering protein sequence data and predicting functional residues are key approaches in protein domain analysis.
- Mapping functional residues to protein structures reveals how substrate specificity variations are encoded structurally.
Purpose of the Study:
- To develop an advanced clustering framework for discovering functional protein domain subfamilies.
- To predict residues critical for determining substrate specificity using a novel computational approach.
- To demonstrate the method's utility across diverse protein domain families.
Main Methods:
- Extension of the context-specific independence mixture model clustering framework.
- Implementation of a partially-supervised learning approach to integrate limited experimental data.
- Application to four protein domain families: phosphatases, pyridoxal dependent decarboxylases, WW, and SH3 domains.
Main Results:
- Discovery of biologically meaningful subfamilies within heterogeneous protein domains.
- Accurate prediction of functional residues linked to substrate specificity.
- Demonstration of the algorithm's effectiveness across multiple protein domain families.
Conclusions:
- The partially-supervised clustering method successfully identifies functional protein domain subfamilies.
- Predicted functional residues offer insights into the structural basis of varying substrate specificities.
- The approach is valuable for analyzing complex protein domain families.
More Related Videos
Related Concept Videos
Protein Families
Protein-protein Interfaces
Ligand Binding and Linkage
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

