Related Experiment Video
Updated: Jun 20, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Enhanced protein fold recognition through a novel data integration approach.
Yiming Ying1, Kaizhu Huang, Colin Campbell
1Department of Engineering Mathematics, University of Bristol, Bristol, BS8 1TR, UK. mathying@gmail.com
This study introduces novel kernel-based methods (MKLdiv-dc and MKLdiv-conv) for integrating diverse data sources in protein fold recognition. These methods achieve state-of-the-art accuracy, improving fold discrimination by over 5%.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Protein fold recognition is crucial for determining 3D protein structures.
- Multiple data sources, including sequence alignments and structural properties, are used for fold discrimination.
- Combining these diverse data sources efficiently is a significant challenge in protein classification.
Purpose of the Study:
- To develop a novel kernel-based approach for integrating multiple heterogeneous data sources for protein fold recognition.
- To propose information-theoretic methods based on Kullback-Leibler (KL) divergence for data integration.
- To evaluate the effectiveness of these methods in improving protein fold classification accuracy.
Main Methods:
- Utilized a kernel-based approach integrating multiple data sources.
- Developed two formulations, MKLdiv-dc and MKLdiv-conv, based on KL divergence between input and output kernel matrices.
- Employed difference of convex (DC) programming for MKLdiv-dc and projected gradient descent for MKLdiv-conv.
Main Results:
- Achieved state-of-the-art performance on the SCOP PDB-40D benchmark dataset for protein fold prediction.
- MKLdiv-dc improved fold discrimination accuracy to 75.19%, a >5% increase over existing methods.
- Demonstrated competitive performance on a yeast protein function prediction task.
Conclusions:
- The proposed MKLdiv-dc and MKLdiv-conv methods effectively integrate diverse data sources for protein fold recognition.
- These methods offer insights into the relative importance of different data sources.
- Achieved superior performance in protein fold prediction and competitive results in protein function prediction.
Related Concept Videos
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Folding
Protein Folding
Protein-protein Interfaces
Protein-Protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...

