Related Experiment Video
Updated: Aug 6, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Local multiple alignment of numerical sequences: detection of subtle motifs from protein sequences and structures
1Bioinformatics Center, Institute for Chemical Research, Kyoto University, Gokasho, Uji, Kyoto 611-0011, Japan. takutsu@kuicr.kyoto-u.ac.jp
Genome Informatics. International Conference on Genome Informatics
|January 16, 2002
Summary
This study introduces a novel computational method for identifying protein motifs using both sequence and structural data. The approach enhances motif discovery, particularly when structural information is integrated.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Biology
Background:
- Identifying conserved motifs in proteins is crucial for understanding protein function and evolution.
- Existing methods often rely solely on sequence data, potentially missing structurally conserved but sequence-divergent motifs.
Purpose of the Study:
- To develop and validate a new computational method for motif discovery integrating protein sequence and structure information.
- To improve the accuracy and scope of motif identification compared to sequence-only approaches.
Main Methods:
- A two-part method involving quantification of protein sequences and structures into numerical representations.
- Development of a novel Gibbs sampling algorithm adapted for real number/vector sequences to find common regions.
- Local multiple alignment to identify fixed-length regions with similar shapes.
Main Results:
- The new method successfully identifies motifs from both multiple protein sequences and structures.
- The integrated approach demonstrates particular utility when structural information is available.
- Comparison with a standard Gibbs sampling program highlights the advantages of the new method.
Conclusions:
- The proposed method offers a powerful new tool for motif discovery in bioinformatics.
- Integrating structural data significantly enhances the identification of conserved protein motifs.
- This approach has implications for functional annotation and drug discovery.
Related Concept Videos
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Signal Sequences and Sorting Receptors
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...

