Some remarks on protein attribute prediction and pseudo amino acid composition
1Gordon Life Science Institute, 13784 Torrey Del Mar Drive, San Diego, CA 92130, USA. kcchou@gordonlifescience.org
Journal of Theoretical Biology
|December 21, 2010
Summary
Computational methods are crucial for identifying protein attributes from sequences, bridging the gap between known sequences and unknown functions. This review focuses on pseudo amino acid composition (PseAAC) for analyzing uncharacterized proteins.
Area of Science:
- * Bioinformatics
- * Computational Biology
- * Proteomics
Background:
- * The rapid increase in protein sequences from genome sequencing outpaces functional characterization.
- * A significant gap exists between sequence-known and attribute-known proteins, hindering research and drug development.
- * Computational methods are needed to rapidly identify protein attributes solely from sequence information.
Purpose of the Study:
- * To review computational methods for identifying protein attributes from sequence data.
- * To discuss key considerations in developing these methods, including datasets, algorithms, accuracy, and web servers.
- * To highlight the role and development of pseudo amino acid composition (PseAAC) in protein attribute prediction.
Main Methods:
- * Review of established computational approaches for protein attribute prediction.
- * Discussion of benchmark dataset construction and protein sample formulation.
- * Focus on pseudo amino acid composition (PseAAC) and its general formulation for feature extraction.
Main Results:
- * Several computational methods have been developed to bridge the sequence-attribute gap.
- * Pseudo amino acid composition (PseAAC) is a key technique for capturing hidden protein sequence features.
- * The general formulation of PseAAC offers flexibility in reflecting complex protein characteristics.
Conclusions:
- * Computational tools are essential for annotating uncharacterized proteins.
- * PseAAC provides a powerful framework for leveraging sequence information to predict protein attributes.
- * Further development of PseAAC can enhance our understanding and utilization of newly discovered proteins.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
What are Proteins?
Proteins are polymers of amino acids linked together by peptide bonds. Proteins and polypeptides are interchangeably used to refer to long chains of amino acids. However, polypeptides have a molecular weight of fewer than 10,000 daltons, while proteins have greater molecular weight. Polypeptides with less than 20 amino acids are called oligopeptides or simply peptides. Interactions among the constituent amino acid side chains of proteins help them fold into a stable 3-dimensional structure...
What are Proteins?
Overview
Amino acids
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible for...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.


