Related Experiment Video
Updated: Feb 14, 2026

Amplification, Next-generation Sequencing, and Genomic DNA Mapping of Retroviral Integration Sites
Published on: March 22, 2016
EnsemGlyPred: Intelligent prediction system for lysine glycation sites integrating deep semantic features and
1School of Artificial Intelligence and Computer Science, Jiangnan University and Engineering Research Center of Intelligent Technology for Healthcare, Ministry of Education, Wuxi, 214122, China.
Motivation:
Protein non-enzymatic glycation plays a crucial role in chronic diseases such as diabetes, atherosclerosis, and neurodegenerative diseases. Accurate identification of lysine glycation sites is essential for elucidating disease pathogenesis and developing diagnostic biomarkers. However, current computational methods face challenges in capturing complex contextual features, integrating multi-source heterogeneous features, and providing biological interpretability.
Results:
This study proposes EnsemGlyPred, an intelligent prediction system integrating multi-dimensional features with weighted ensemble learning. We constructed a reliable benchmark dataset from the PLMD database after rigorous data processing. A multi-level feature extraction framework was designed, integrating AAC features (amino acid composition), PAAC features (incorporating sequence order), and deep semantic features from ProGen2. Three optimized base classifiers were constructed: AAC-based XGBoost, PAAC-based XGBoost, and ProGen2-based BiLSTM models, with scientifically allocated weights (0.33, 0.4, 0.27) for ensemble prediction. Rigorous ten-fold cross-validation and independent testing demonstrated that the ensemble model significantly outperformed existing methods, particularly achieving significant improvement in recall metrics. Systematic ablation experiments confirmed the effectiveness of multi-feature fusion and weighted ensemble strategies. Feature importance analysis revealed the key regulatory role of amino acids in the lysine neighboring region and positively charged residues. t-SNE visualization demonstrated improved discriminative feature space distribution. This research provides a systematic methodological framework for computational prediction of protein post-translational modifications, offering reliable theoretical foundation and important guidance for related fields.
Availability:
an intelligent online prediction platform (http://www.ensemglypred.com) has been developed to facilitate the efficient conduct of academic research. The predictive model EnsemGlyPred and associated datasets are available at: https://github.com/pigyoo-1009/EnsemGlyPred.
More Related Videos
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conserved Binding Sites
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Ligand Binding Sites
Intelligence
Predicting Molecular Geometry