ACP-ADA: A Boosting Method with Data Augmentation for Improved Prediction of Anticancer Peptides
Sadik Bhattarai1, Kyu-Sik Kim2, Hilal Tayara3
1Department of Electronics and Information Engineering, Jeonbuk National University, Jeonju 54896, Korea.
Abstract:
Cancer is the second-leading cause of death worldwide, and therapeutic peptides that target and destroy cancer cells have received a great deal of interest in recent years. Traditional wet experiments are expensive and inefficient for identifying novel anticancer peptides; therefore, the development of an effective computational approach is essential to recognize ACP candidates before experimental methods are used. In this study, we proposed an Ada-boosting algorithm with the base learner random forest called ACP-ADA, which integrates binary profile feature, amino acid index, and amino acid composition with a 210-dimensional feature space vector to represent the peptides. Training samples in the feature space were augmented to increase the sample size and further improve the performance of the model in the case of insufficient samples. Furthermore, we used five-fold cross-validation to find model parameters, and the cross-validation results showed that ACP-ADA outperforms existing methods for this feature combination with data augmentation in terms of performance metrics. Specifically, ACP-ADA recorded an average accuracy of 86.4% and a Mathew's correlation coefficient of 74.01% for dataset ACP740 and 90.83% and 81.65% for dataset ACP240; consequently, it can be a very useful tool in drug development and biomedical research.
Insights
Identifying anticancer peptides (ACPs) computationally is crucial for drug development. A new Ada-boosting algorithm, ACP-ADA, effectively predicts ACP candidates using integrated features and data augmentation, outperforming existing methods.
Area of Science:
- Biotechnology
- Computational Biology
- Oncology
Background:
- Cancer is a leading global cause of death.
- Therapeutic peptides targeting cancer cells are of significant interest.
- Identifying novel anticancer peptides (ACPs) via traditional experiments is costly and inefficient.
Purpose of the Study:
- To develop an effective computational approach for identifying anticancer peptide candidates.
- To propose a novel machine learning model for ACP prediction.
Main Methods:
- An Ada-boosting algorithm (ACP-ADA) was developed using random forest as the base learner.
- Peptides were represented using a 210-dimensional feature vector integrating binary profile, amino acid index, and amino acid composition.
- Training samples were augmented to enhance model performance with limited data.
Main Results:
- ACP-ADA demonstrated superior performance compared to existing methods.
- Five-fold cross-validation was used to optimize model parameters.
- The model achieved high accuracy (e.g., 86.4% on ACP740) and Mathew's correlation coefficient (e.g., 74.01% on ACP740).
Conclusions:
- ACP-ADA is a highly effective computational tool for identifying anticancer peptides.
- The integration of diverse features and data augmentation significantly improves prediction accuracy.
- This approach can accelerate drug development and advance biomedical research.


