Related Experiment Videos
Molecular sequence accuracy and the analysis of protein coding regions
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894.
Summary
Molecular sequence errors impact data interpretation. Simultaneous translation and alignment algorithms enhance protein homology identification, remaining robust even with significant error rates.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Experimental molecular sequence data inherently contain errors.
- The impact of these errors depends on the analytical methods used.
- Understanding error resilience is crucial for accurate biological data interpretation.
Purpose of the Study:
- To investigate the effect of nucleic acid sequence errors on amino acid sequence alignment.
- To assess the robustness of homology identification algorithms against sequence errors.
- To evaluate methods for improving error tolerance in sequence analysis.
Main Methods:
- Utilized a simultaneous translation and alignment algorithm.
- Introduced controlled rates of frameshifting (insertion/deletion) and base substitution errors.
- Incorporated prior knowledge of error characteristics and biased codon usage (Saccharomyces cerevisiae).
Main Results:
- Homology identification is resilient to random errors using the simultaneous algorithm.
- Proteins with >30% sequence identity are reliably recognized despite 1% frameshift and 5% substitution errors.
- Prior knowledge significantly improves error tolerance in alignments and reading frame detection.
Conclusions:
- Simultaneous translation and alignment offers robust protein homology detection.
- Incorporating error-specific information enhances the reliability of sequence analysis.
- Accurate reading frame identification in yeast is achievable even with sequence imperfections.