Related Experiment Video
Updated: Jul 31, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Detection of protein coding sequences using a mixture model for local protein amino acid sequence
E C Thayer1, C Bystroff, D Baker
1Department of Biochemistry, University of Washington, Seattle 98105, USA.
Summary
This study introduces recurrent amino acid patterns to improve gene detection in genomic DNA. These patterns help distinguish coding DNA from non-coding DNA, potentially enhancing current gene-finding tools.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate gene detection is crucial for understanding genomic data from large-scale sequencing.
- Existing gene-finding methods rely on DNA content differences and regulatory site recognition.
Purpose of the Study:
- To investigate the utility of short, recurrent amino acid sequence patterns (3-19 amino acids) as a novel statistic for gene detection.
- To assess if these patterns can improve the accuracy of identifying protein-coding regions.
Main Methods:
- Developed a finite mixture model incorporating recurrent amino acid patterns.
- Used the model to discriminate between protein sequences, randomized sequences, and short non-coding DNA segments from the S. cerevisiea genome.
- Compared scores derived from human exons using this model against existing GENSCAN scores.
Main Results:
- The finite mixture model demonstrated partial discrimination capabilities between coding and non-coding sequences.
- Scores generated by the new model for human exons showed no correlation with GENSCAN scores.
- This lack of correlation suggests complementary functionality.
Conclusions:
- Recurrent amino acid patterns offer a valuable content statistic for gene finding.
- Integrating this protein pattern recognition module into current gene recognition programs could enhance their performance.
- This approach provides a new avenue for improving the accuracy of identifying protein-coding regions in genomic DNA.
Related Concept Videos
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Leaky Scanning
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Tagging and Fusion Proteins
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
Peptide Identification Using Tandem Mass Spectrometry
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Proteomics
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...

