IDPpred: a new sequence-based predictor for identification of intrinsically disordered protein with enhanced accuracy

Deepak Chaurasiya1, Rajkrishna Mondal2, Tapobrata Lahiri1

  • 1Department of Applied Sciences, Indian Institute of Information Technology, Prayagraj, UP, India.

Discovery of intrinsically disordered proteins (IDPs) and protein hybrids that contain both intrinsically disordered protein regions (IDPRs) along with ordered regions has changed the sequence-structure-function paradigm of protein. These proteins with lack of persistently fixed structure are often found in all organisms and play vital roles in various biological processes. Some of them are considered as potential drug targets due to their overrepresentation in pathophysiological processes. The major bottlenecks for characterizing such proteins are their occasional overexpression, difficulty in getting purified homogeneous form and the challenge of investigating them experimentally. Sequence-based prediction of intrinsic disorder remains a useful strategy especially for many large-scale proteomic investigations. However, worst accuracy still occurs for short disordered regions with less than ten residues, for the residues close to order-disorder boundaries, for regions that undergo coupled folding and binding in presence of partner, and for prediction of fully disordered proteins. Annotation of fully disordered proteins mostly relies on the far-UV circular dichroism experiment which gives overall secondary structure composition without residue-level resolution. Current methods including that using secondary structure information failed to predict half of target IDPs correctly in the recent Critical Assessment of protein Intrinsic Disorder prediction (CAID) experiment. This study utilized profiles of random sequential appearance of physicochemical properties of amino acids and random sequential appearance of order and disorder promoting amino acids in protein together with the existing CIDER feature for the prediction of IDP from sequence input. Our method was found to significantly outperform the existing predictors across different datasets.Communicated by Ramaswamy H. Sarma.

Related Concept Videos

Intrinsically Disordered Proteins02:18

Intrinsically Disordered Proteins

Intrinsically disordered proteins are a group of proteins that do not fold into specific three-dimensional structures. Their structural flexibility allows them to complement ordered proteins to perform functions that are inaccessible to rigid structures. They are more common in eukaryotes than prokaryotes and may either be exclusively intrinsically disordered or hybrid proteins, consisting of a mix of ordered and disordered regions. The absence of a rigid structure in these proteins can be...
17.9K
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein-protein Interfaces02:04

Protein-protein Interfaces

Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
15.4K
Signal Sequences and Sorting Receptors01:41

Signal Sequences and Sorting Receptors

Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.4K
Mitochondrial Precursor Proteins01:39

Mitochondrial Precursor Proteins

Mitochondrial precursors are partially unfolded or loosely folded polypeptide chains. Newly synthesized precursors are inhibited from spontaneously folding into their native conformation by the cytosolic chaperones, heat shock proteins 70 (Hsp70), and mitochondrial import stimulation factors (MSFs). Precursors bound to MSFs are guided to the TOM70-TOM37 receptors, while precursors bound to Hsp70  chaperones are targetted to TOM20-TOM22 receptor complexes.
Most of the mitochondrial...
2.6K