Computational identification of MoRFs in protein sequences

Nawar Malhis1, Jörg Gsponer2

  • 1Centre for High-Throughput Biology and Department of Biochemistry and Molecular Biology, University of British Columbia, Vancouver, BC V6T 1Z4, Canada.

Abstract

Insights

We developed MoRFCHiBi, a novel computational tool for accurately predicting intrinsically disordered protein regions known as molecular recognition features (MoRFs). This fast and accurate MoRF prediction method outperforms existing tools, aiding in understanding protein regulation.

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Protein Structure Prediction

Background:

  • Intrinsically disordered protein regions are crucial for biological regulation.
  • Molecular recognition features (MoRFs) drive disorder-to-order transitions upon binding.
  • Accurate prediction of MoRFs in protein sequences is a significant computational challenge.

Purpose of the Study:

  • To introduce MoRFCHiBi, a novel computational approach for the fast and accurate prediction of MoRFs.
  • To improve upon existing methods for identifying MoRFs in protein sequences.

Main Methods:

  • MoRFCHiBi employs a hybrid approach combining two support vector machine (SVM) models.
  • The models utilize distinct kernels to leverage amino acid composition differences and sequence similarities.
  • High noise tolerance is achieved through the chosen SVM kernels.

Main Results:

  • MoRFCHiBi demonstrated superior performance compared to existing MoRF predictors (MoRFpred and ANCHOR) on established test datasets.
  • The predictor achieved higher accuracy across various evaluation metrics.
  • MoRFCHiBi is computationally fast and available for download, facilitating its integration into other prediction pipelines.

Conclusions:

  • MoRFCHiBi represents a significant advancement in MoRF prediction accuracy and speed.
  • Its performance and accessibility make it a valuable tool for researchers studying protein regulation and intrinsically disordered proteins.

Related Concept Videos

Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
15.2K
Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
17.7K
Protein Families02:47

Protein Families

4.7K
Conservation of Protein Domains02:26

Conservation of Protein Domains

4.4K
Protein Networks02:26

Protein Networks

An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.7K
Globular and Fibrous Proteins02:21

Globular and Fibrous Proteins

Many proteins can be classified into two distinct subtypes - globular or fibrous. These two types differ in their shapes and solubilities.
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
49.3K