DeepRAGIL-2: a retrieval-augmented protein language model framework for sensitive and accurate prediction of
Juan Peter Timothy Yuune1, Van The Le1, Yu-Yen Ou1,2
1Department of Computer Science and Engineering, Yuan Ze University, Taoyuan City, Taiwan.
Abstract:
Interleukin-2 (IL-2) is a pleiotropic cytokine central to T-cell activation, proliferation, and immunological tolerance, yet existing computational tools for predicting IL-2-inducing peptides suffer from near-zero sensitivity under realistic class-imbalanced conditions, rendering them impractical for positive-class discovery. Here we present DeepRAGIL-2, a sequence-based framework integrating ESM-2 protein language model embeddings, dual-window convolutional feature extraction, and retrieval-augmented embedding fusion (RAG) into a unified classification pipeline. Trained on 3,825 redundancy-filtered experimentally validated sequences from IEDB and evaluated on an independent test set under a 1:4 class imbalance, DeepRAGIL-2 achieves a sensitivity of 0.7501, specificity of 0.9708, MCC of 0.7605, and AUC of 0.8654, outperforming conventional machine learning classifiers and published IL-2 prediction tools on sensitivity and MCC. Across ten independent runs, performance was stable at MCC of 0.7555 ± 0.0173, confirming that the reported result is representative rather than an artefact of a single weight initialisation. Ablation experiments identify RAG integration as the single most consequential architectural component, with MCC rising from 0.3453 to 0.7605 upon its inclusion, demonstrating that retrieval of experimentally validated reference sequences at inference time substantially compensates for the limited positive-class training data available for this task. DeepRAGIL-2 represents the first IL-2 prediction framework with demonstrated sensitivity for positive-class recovery under realistic conditions, establishing a framework applicable in principle to cytokine-specific immunological sequence classification.


