Related Experiment Video
Updated: Dec 18, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
69.6K
ProtPCV: A Fixed Dimensional Numerical Representation of Protein Sequence to Significantly Reduce Sequence Search
Manoj Kumar Pal1, Tapobrata Lahiri2, Rajnish Kumar1
1Department of Applied Sciences, Indian Institute of Information Technology, Allahabad, UP, 211015, India.
Interdisciplinary Sciences, Computational Life Sciences
|June 12, 2020
Summary
This study introduces a novel numerical representation for protein sequences, enabling faster homologue identification. This method enhances protein similarity searches and reduces computational complexity.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Science
Background:
- Current protein homologue search methods often overlook the ordered physicochemical properties of protein sequences.
- Existing techniques lack robust benchmarking, questioning their reliability in search engines.
Purpose of the Study:
- To develop a novel numerical representation for protein sequences to improve homologue identification.
- To establish a more comprehensive benchmarking approach for protein similarity search methods.
Main Methods:
- A fixed-dimensional numerical representation of protein sequences was created, extending the periodicity count concept.
- Euclidean distance was employed as a direct similarity measure between protein representations.
- The new method was benchmarked against BLAST, PSI-BLAST, Needleman-Wunsch, and Smith-Waterman using novel correlation-based metrics.
Main Results:
- The novel numerical representation significantly reduces computational complexity in protein sequence searches to O(log(n)).
- The proposed benchmarking methods provide a stronger validation of search technique performance.
- The approach demonstrates potential for various similarity-based operations like clustering and phylogenetic analysis.
Conclusions:
- This new numerical representation offers an efficient and validated method for protein sequence analysis.
- The findings suggest improved accuracy and speed for protein homologue detection and related applications.
- The method facilitates broader applications in protein classification and evolutionary studies.
Related Concept Videos
Protein Families
16.5K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.5K
Conserved Binding Sites
4.9K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.9K
Protein Organization
8.7K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
8.7K
Protein Organization
155.2K
Overview
155.2K
Conservation of Protein Domains Over Different Proteins
13.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.9K
Signal Sequences and Sorting Receptors
14.1K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
14.1K

