Related Experiment Video
Updated: Jun 19, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
A stochastic context free grammar based framework for analysis of protein sequences
Witold Dyrka1, Jean-Christophe Nebel
1Institute of Biomedical Engineering and Instrumentation, Wroclaw University of Technology, Poland. Witold.Dyrka@pwr.wroc.pl
We developed a new Stochastic Context Free Grammar framework to analyze protein sequences and identify binding sites. This method produces human-readable descriptors, improving upon current machine learning techniques for protein analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Proteomics
Background:
- Formal language theory has seen limited application in proteomics due to the protein alphabet size and amino acid complexity.
- Existing methods struggle with higher-order dependencies like nested and crossing relationships in proteins.
- Stochastic regular grammars lack the expressive power for complex protein structures.
Purpose of the Study:
- To introduce a Stochastic Context Free Grammar (SCFG) based framework for protein sequence analysis.
- To develop a system for producing binding site descriptors that offer insights into protein structure.
- To overcome limitations of existing methods in handling complex protein dependencies.
Main Methods:
- Induced grammars using a genetic algorithm for protein sequence analysis.
- Utilized quantitative amino acid properties to manage the protein alphabet size.
- Applied structural constraints to grammars to optimize the rule search space.
- Combined grammars based on different properties for comprehensive information capture.
Main Results:
- Developed a system producing human-readable binding site descriptors.
- Achieved high accuracy in both annotation and detection of protein binding sites.
- Demonstrated suitability for patterns shared by non-homologous proteins, outperforming current methods.
- Descriptors highlight biologically meaningful features and provide structural insights.
Conclusions:
- Introduced a novel SCFG framework for generating binding site descriptors in protein sequence analysis.
- Validated the framework's effectiveness, producing human-readable descriptors.
- The approach surpasses current machine learning techniques in analyzing complex binding sites.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Genome Annotation and Assembly
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme can...
Protein and Protein Structures
A protein's shape is critical to its function. For example, an enzyme can...
Protein Folding Quality Check in the RER

