Detection of new protein domains using co-occurrence: application to Plasmodium falciparum

Nicolas Terrapon1, Olivier Gascuel, Eric Maréchal

  • 1Méthodes et algorithmes pour la Bioinformatique, LIRMM, Université Montpellier 2, CNRS, 161 rue Ada, 34392 Montpellier Cedex 5, France.

Abstract

Insights

This study introduces a new method to enhance protein domain detection in Plasmodium falciparum, identifying hundreds of novel Pfam domains and improving Gene Ontology annotations for malaria research.

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Hidden Markov Models (HMMs) are effective for protein domain identification.
  • Highly divergent proteins, such as those in Plasmodium falciparum, often lead to missed domain detections.
  • Plasmodium falciparum is the primary cause of human malaria.

Purpose of the Study:

  • To improve the sensitivity of HMM-based protein domain detection.
  • To identify novel protein domains in Plasmodium falciparum.
  • To enhance functional annotation of the Plasmodium proteome.

Main Methods:

  • Developed a novel method leveraging the co-occurrence patterns of protein domains.
  • Utilized existing Pfam and InterPro domain information to infer presence of missed domains.
  • Employed a shuffling procedure to estimate the false discovery rate.

Main Results:

  • Identified 585 new Pfam domains in Plasmodium falciparum, adding to the 3683 known domains.
  • Achieved an estimated false discovery rate below 20% for the newly identified domains.
  • Generated 387 new Gene Ontology (GO) annotations for the P. falciparum proteome.
  • Obtained similar results for related Plasmodium species (P. vivax, P. yoelii).

Conclusions:

  • The proposed method significantly enhances protein domain detection sensitivity for divergent proteins.
  • The newly identified domains and GO annotations provide valuable insights into Plasmodium biology.
  • This approach aids in understanding malaria parasite genetics and developing new interventions.

Related Concept Videos

Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains02:26

Conservation of Protein Domains

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key locations, protein...