Related Experiment Video
Updated: Aug 6, 2026

05:53
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
Two methods for improving performance of an HMM and their application for gene finding
1Center for Biological Sequence Analysis, Technical University of Denmark, Lyngby, Denmark. krogh@cbs.dtu.dk
Summary
This study introduces a novel approach to training hidden Markov models (HMMs) for gene finding. By optimizing for correct predictions rather than just sequence probability, the HMM-gene finder achieves significantly improved performance.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Hidden Markov models (HMMs) are widely used for gene finding, incorporating submodels for various genomic regions.
- Traditional methods often estimate HMM components independently, which may not be optimal for the overall model performance.
- The standard maximum likelihood estimation (MLE) criterion focuses on maximizing DNA sequence probability, not prediction accuracy.
Purpose of the Study:
- To develop an improved method for estimating HMMs for gene finding.
- To introduce a new training criterion that maximizes the probability of correct predictions.
- To enhance the performance of gene-finding algorithms like HMM-gene.
Main Methods:
- Estimating the entire HMM as a whole from labeled sequences, rather than independent sub-sequences.
- Utilizing a conditional maximum likelihood (CML) criterion to maximize prediction accuracy.
- Developing an approximate algorithm to find the most probable prediction summed over all paths yielding the same prediction.
Main Results:
- The proposed CML criterion is more effective for training gene-finding HMMs than standard MLE.
- The new approximate algorithm efficiently computes the most probable prediction.
- These methods significantly contribute to the high performance of the HMM-gene finder.
Conclusions:
- Training HMMs using conditional maximum likelihood enhances gene-finding accuracy.
- The developed algorithm and training strategy lead to substantial performance improvements in gene prediction.
- This approach offers a more effective way to build and train HMMs for genomic analysis.
Related Concept Videos
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
Pharmacogenomics: Identification of New Drug Targets
Advances in genomics have profoundly influenced drug discovery by increasing both the speed and accuracy of pharmaceutical development. Pharmacogenomics, which examines how genetic variation influences drug response, facilitates the identification of novel therapeutic targets and enables patient stratification for personalized treatment. These strategies contribute to improved drug efficacy, minimized adverse effects, and more efficient clinical trial design.Mapping genetic differences...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...

