DextMP: deep dive into text for predicting moonlighting proteins

Ishita K Khan1, Mansurul Bhuiyan2, Daisuke Kihara1,3

  • 1Department of Computer Science, Purdue University, West Lafayette, IN, USA.

Abstract

Insights

Moonlighting proteins (MPs), which have multiple functions, are crucial in biology and disease. A new method, DextMP, accurately identifies MPs using text analysis, revealing a significant percentage of potential MPs across species.

Area of Science:

  • Proteomics
  • Bioinformatics
  • Computational Biology

Background:

  • Moonlighting proteins (MPs) perform multiple distinct cellular functions, impacting biological systems and disease.
  • Current biological databases lack specific labeling for MPs, hindering accurate functional annotation.
  • Understanding MPs is critical for advancing computational function prediction and database annotation.

Purpose of the Study:

  • To develop a novel computational method, DextMP, for predicting moonlighting proteins.
  • To leverage textual features from scientific literature and the UniProt database for MP prediction.

Main Methods:

  • DextMP extracts textual information including titles, abstracts, and UniProt function descriptions.
  • Compares three language models: deep unsupervised learning, Term Frequency-Inverse Document Frequency (TF-IDF), and Latent Dirichlet Allocation (LDA).
  • Utilizes cross-validation on known MP and non-MP datasets for performance evaluation.

Main Results:

  • DextMP achieved over 91% accuracy in predicting moonlighting proteins, outperforming existing methods.
  • Analysis of human, yeast, and Xenopus laevis genomes identified 2.5-35% of proteomes as potential MPs.
  • The study highlights the prevalence and significance of moonlighting proteins across different species.

Conclusions:

  • DextMP provides an accurate and effective approach for identifying moonlighting proteins.
  • The findings suggest a substantial proportion of proteomes consist of moonlighting proteins, necessitating further research.
  • This work enhances the computational prediction and annotation of protein functions.