Related Experiment Video
Updated: Jul 16, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
HMM-ModE--improved classification using profile hidden Markov models by optimising the discrimination threshold and
Prashant K Srivastava1, Dhwani K Desai, Soumyadeep Nandi
1School of Information Technology, Jawaharlal Nehru University, New Delhi, India. prashant.k.srivastava@gmail.com <prashant.k.srivastava@gmail.com>
This study introduces HMM-ModE, a protocol that enhances protein family classification using profile Hidden Markov Models (HMMs). By incorporating negative training data, HMM-ModE significantly improves the specificity of identifying proteins based on their molecular function.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Profile Hidden Markov Models (HMMs) are statistical tools for protein family identification based on conserved sequence patterns.
- Existing HMMs can be improved by using pre-classified negative training sequences to enhance specificity.
- Fold-specific and function-specific signals influence HMM accuracy.
Purpose of the Study:
- To develop and validate a protocol (HMM-ModE) for generating family-specific HMMs with improved specificity.
- To optimize HMMs by minimizing the influence of non-discriminating signals and enhancing function-specific signals.
- To leverage negative training data for more accurate protein classification.
Main Methods:
- Constructing an initial profile HMM from a protein family's aligned sequences.
- Utilizing negative training sequences to identify and correct false positives by optimizing threshold cutoffs.
- Modifying HMM emission probabilities based on alignments of true and false positive sequences.
- Employing ten-fold cross-validation for threshold optimization.
Main Results:
- The HMM-ModE protocol significantly improved specificity in classifying AGC kinase sub-families from an average of 21% to 98%.
- Further refinement by modifying model probabilities increased average specificity to 99% for kinase sub-families.
- Similar specificity enhancements were observed for G-Protein coupled receptors with lower sequence identity (20.6%).
- The protocol was successfully applied to high-throughput classification of protein kinases.
Conclusions:
- HMM-ModE effectively maximizes the contribution of discriminating residues for protein classification based on molecular function.
- The protocol demonstrates high specificity and potential for widespread application in sequence annotation.
- The method's utility is enhanced by the increasing availability of pre-classified sequence data.