Related Experiment Videos
Remote homology detection based on oligomer distances
Thomas Lingner1, Peter Meinicke
1Abteilung Bioinformatik, Institut für Mikrobiologie und Genetik, Georg-August-Universität Göttingen Goldschmidtstr. 1, 37077 Göttingen, Germany. thomas@gobics.de
Bioinformatics (Oxford, England)
|July 14, 2006
Summary
A new feature vector representation for protein sequences based on oligomer distances offers a faster and more interpretable alternative to traditional kernel methods for remote homology detection.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Sequence Analysis
Background:
- Remote homology detection is a critical challenge in bioinformatics.
- Kernel-based methods are accurate but computationally expensive and lack interpretability.
- Hyperparameters in existing methods complicate cross-dataset applications.
Purpose of the Study:
- To develop a novel, efficient, and interpretable method for remote homology detection.
- To overcome the limitations of current kernel-based approaches.
Main Methods:
- Introduced a feature vector representation for protein sequences.
- Utilized distances between short oligomers (K-mers) to construct feature spaces.
- Employed distance histograms for pairs of K-mers.
Main Results:
- The new distance-based approach significantly improves computational speed.
- Achieved highly competitive prediction performance compared to state-of-the-art methods.
- The resulting model allows for easy analysis of discriminative features.
- Eliminated the need for kernel hyperparameter tuning.
Conclusions:
- This novel representation offers a computationally efficient and interpretable solution for remote homology detection.
- The method is competitive with existing approaches and simplifies application across diverse datasets.