Related Experiment Videos
A hidden Markov model that finds genes in E. coli DNA
A Krogh1, I S Mian, D Haussler
1Nordita, Copenhagen, Denmark.
Nucleic Acids Research
|November 11, 1994
Summary
A novel hidden Markov model (HMM) accurately identifies protein-coding genes in E. coli DNA. This computational tool aids in discovering new genes and potential sequencing errors within the E. coli genome.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying protein-coding genes is crucial for understanding genome function.
- Existing methods may face challenges with raw genomic data containing errors or frameshifts.
Purpose of the Study:
- To develop and validate a hidden Markov model (HMM) for precise gene prediction in E. coli.
- To improve the detection of known and potentially novel genes in E. coli DNA sequences.
Main Methods:
- Developed a hidden Markov model (HMM) incorporating codon frequencies and intergenic region patterns specific to E. coli.
- Modeled potential sequencing errors like insertions and deletions within codons.
- Estimated HMM parameters using ~1 million nucleotides and tested on ~325,000 nucleotides of E. coli genome data.
Main Results:
- The HMM successfully located the exact positions of approximately 80% of known E. coli genes.
- It also provided approximate locations for an additional 10% of known genes.
- Identified several potential novel genes and flagged regions with possible sequencing errors or frameshifts.
Conclusions:
- The developed HMM is an effective tool for E. coli gene prediction.
- The model demonstrates utility in discovering new genes and identifying data quality issues in genomic sequences.