Related Experiment Video
Updated: Jul 4, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Word correlation matrices for protein sequence analysis and remote homology detection
Thomas Lingner1, Peter Meinicke
1Department of Bioinformatics, Institute of Microbiology and Genetics, Georg-August-University Göttingen, Göttingen, Germany. thomas@gobics.de
A new protein sequence classification method uses average word similarity for fast and interpretable predictions. This approach achieves competitive performance in protein remote homology detection and identifies characteristic sequence regions.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Protein sequence classification is crucial in computational biology.
- Kernel-based methods offer high accuracy but lack interpretability and are computationally expensive.
- Interpretable models are needed for analyzing discriminative sequence features.
Purpose of the Study:
- To develop a novel, interpretable, and computationally efficient kernel for protein sequence classification.
- To enable fast classification of new protein sequences and analysis of discriminative features.
- To improve protein remote homology detection.
Main Methods:
- A novel kernel based on average word similarity between protein sequences was developed.
- The kernel creates a feature space allowing for analysis of discriminative features.
- The method was evaluated on a benchmark dataset for protein remote homology detection.
Main Results:
- The proposed word correlation approach achieves highly competitive performance compared to state-of-the-art methods.
- The method provides an interpretable model in terms of biologically meaningful features.
- Analysis of discriminative words identifies characteristic regions in biological sequences.
Conclusions:
- The novel word correlation kernel offers a powerful tool for protein remote homology detection.
- Its interpretability facilitates the identification of key sequence features.
- High computational efficiency allows application to large-scale database searches for potential homologs.
More Related Videos
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes

