Related Experiment Video
Updated: Jun 16, 2026

12:27
Measuring Peptide Translocation into Large Unilamellar Vesicles
Published on: January 27, 2012
13.9K
Cell-penetrating peptides predictors: A comparative analysis of methods and datasets
Karen Guerrero-Vázquez1,2, Gabriel Del Rio3, Carlos A Brizuela1
1Department of Computer Science, CICESE Research Center, Ensenada, 22860, Mexico.
Molecular Informatics
|September 6, 2023
Summary
Dataset similarity significantly impacts Cell-Penetrating Peptide (CPP) predictor performance. High sequence similarity in training data makes prediction challenging, highlighting the need for robust datasets to evaluate new CPP predictors effectively.
Area of Science:
- Biochemistry
- Bioinformatics
- Drug Discovery
Background:
- Cell-Penetrating Peptides (CPP) offer therapeutic potential beyond small molecules.
- Numerous predictors exist for identifying and designing novel CPPs.
- Previous performance comparisons of CPP predictors lacked clarity on dataset influence.
Purpose of the Study:
- To systematically evaluate the impact of peptide sequence similarity within datasets on CPP predictor performance.
- To determine whether dataset characteristics, model choice, or descriptors are primary drivers of predictor accuracy.
Main Methods:
- Analysis of CPP predictor performance across datasets with varying degrees of sequence similarity.
- Comparison of classifier outcomes based on the sequence relatedness of positive (CPP) and negative (non-CPP) examples.
Main Results:
- Dataset characteristics, specifically sequence similarity, exert a greater influence on predictor performance than the chosen model or descriptors.
- Classifiers perform well on datasets with low sequence similarity between CPP and non-CPP examples.
- Datasets with high sequence similarity between CPP and non-CPP examples present a significant challenge for accurate prediction.
Conclusions:
- The composition and sequence similarity of training datasets are critical factors in assessing CPP predictor reliability.
- Future evaluations of new CPP predictors should utilize datasets with high sequence similarity to ensure rigorous performance assessment.
- Understanding dataset influence is key to developing more accurate and generalizable CPP identification tools.

