ACP-ADA: A Boosting Method with Data Augmentation for Improved Prediction of Anticancer Peptides

Sadik Bhattarai1, Kyu-Sik Kim2, Hilal Tayara3

  • 1Department of Electronics and Information Engineering, Jeonbuk National University, Jeonju 54896, Korea.

Insights

Identifying anticancer peptides (ACPs) computationally is crucial for drug development. A new Ada-boosting algorithm, ACP-ADA, effectively predicts ACP candidates using integrated features and data augmentation, outperforming existing methods.

Area of Science:

  • Biotechnology
  • Computational Biology
  • Oncology

Background:

  • Cancer is a leading global cause of death.
  • Therapeutic peptides targeting cancer cells are of significant interest.
  • Identifying novel anticancer peptides (ACPs) via traditional experiments is costly and inefficient.

Purpose of the Study:

  • To develop an effective computational approach for identifying anticancer peptide candidates.
  • To propose a novel machine learning model for ACP prediction.

Main Methods:

  • An Ada-boosting algorithm (ACP-ADA) was developed using random forest as the base learner.
  • Peptides were represented using a 210-dimensional feature vector integrating binary profile, amino acid index, and amino acid composition.
  • Training samples were augmented to enhance model performance with limited data.

Main Results:

  • ACP-ADA demonstrated superior performance compared to existing methods.
  • Five-fold cross-validation was used to optimize model parameters.
  • The model achieved high accuracy (e.g., 86.4% on ACP740) and Mathew's correlation coefficient (e.g., 74.01% on ACP740).

Conclusions:

  • ACP-ADA is a highly effective computational tool for identifying anticancer peptides.
  • The integration of diverse features and data augmentation significantly improves prediction accuracy.
  • This approach can accelerate drug development and advance biomedical research.