Related Experiment Video
Updated: Sep 16, 2026

Adoptive Immunotherapy of iNKT Cells in Glucose-6-Phosphate Isomerase (G6PI)-Induced RA Mice
Published on: January 31, 2020
GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13-Inducing Peptides
Hiroyuki Kurata1, Hiroto Tsuruta1, Soyogu Shigetomi1
1Department of Bioscience and Bioinformatics, Kyushu Institute of Technology, 680-4 Kawazu, Iizuka 820-8502, Fukuoka, Japan.
Abstract:
Identifying interleukin-6 (IL-6) and interleukin-13 (IL-13)-inducing peptides is important for drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening is costly and time-consuming. Machine learning and deep learning models have been developed that distinguish functional peptides from no-function ones, but their performance is limited by the small number of experimentally validated peptides. In this study, we propose a generative AI-driven data augmentation framework, GDA, and its prediction system, GDA-Pred, to improve the performance of state-of-the-art (SOTA) classifiers under limited data. GDA generates peptide sequences using three generative models: generative adversarial networks, diffusion models, and variational autoencoders. The framework is controlled by four hyperparameters: generative model type, sequence identity cutoff, probability threshold, and augmentation ratio. Because optimizing these hyperparameters is difficult with small datasets, we used anti-inflammatory peptide (AIP) data as a proof-of-concept to identify an effective reference hyperparameter setting. We evaluated GDA using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out benchmark test. The GDA with the AIP-derived reference hyperparameter setting was then applied to SOTA classifiers to identify IL-6 and IL-13-inducing peptides as a case study. GDA-Pred consistently improved prediction performance for both cytokine-inducing peptide datasets, demonstrating the potential of generative AI to overcome data scarcity in peptide prediction.
