Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data

Katarzyna Sidorczuk1, Przemysław Gagat1, Filip Pietluch1

  • 1University of Wrocław, Faculty of Biotechnology, Poland.

Insights

Negative data sampling significantly impacts antimicrobial peptide (AMP) prediction models. Ensuring consistent data generation for training and benchmarking is crucial for accurate AMP discovery and reliable performance evaluation.

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Drug Discovery

Background:

  • Antimicrobial peptides (AMPs) show promise against microbes, viruses, and cancer cells, offering an alternative to traditional antibiotics due to lower resistance.
  • Machine learning (ML) is a cost-effective approach for discovering novel AMPs, leading to the development of numerous computational prediction tools.
  • The performance of these AMP prediction tools is heavily influenced by the data used for training and benchmarking.

Purpose of the Study:

  • To investigate the impact of negative data sampling methods on the performance of machine learning models for antimicrobial peptide (AMP) prediction.
  • To evaluate the reliability of current benchmarking practices for AMP prediction software.
  • To provide a standardized platform for fair and accurate benchmarking of AMP predictors.

Main Methods:

  • Generated 660 predictive models using 12 distinct ML architectures and 11 different negative data sampling strategies.
  • Defined model architectures and sampling methods based on established AMP prediction software.
  • Developed a web server, AMPBenchmark, for standardized and unbiased performance evaluation.

Main Results:

  • Model performance is positively affected when training and benchmark datasets are generated using the same or similar negative data sampling methods.
  • Current benchmarking analyses for AMP prediction models are significantly biased due to inconsistent data sampling.
  • The true accuracy of existing AMP prediction models remains uncertain.

Conclusions:

  • Inconsistent negative data sampling in training and benchmarking leads to biased evaluations of antimicrobial peptide (AMP) prediction models.
  • A standardized approach to data generation is essential for reliable performance assessment.
  • The AMPBenchmark web server offers a solution for fair and accurate benchmarking, aiding researchers in selecting optimal AMP prediction tools.