Related Experiment Video
Updated: Sep 25, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
GLIMP: An Integrated Graph Neural Network-Large Language Model for Promoter Recognition
Abstract:
Accurate identification of DNA promoter regions is crucial for understanding gene regulation, yet remains challenging. Existing models often struggle to generalize across species or promoter types due to their inability to model functional degeneracy and capture subtle regulatory signals. To address this, we propose GLIMP, a novel promoter classification framework that integrates a large-scale genomic language model with Graph Neural Networks (GNNs). The core innovation of GLIMP lies in its multimodal fusion of a Shapelet-guided Weighted Context-Aware Interaction Network (WC-IDNG) graph representation-which topologically connects functionally similar but sequence-divergent motifs-and deep semantic features extracted by DNABERT. Specifically, we construct a WC-IDNG graph from $k$-mers to capture local proximity and structural similarity, refined by biologically meaningful shapelets. These graph-aggregated signals are then fused with global contextual semantics from DNABERT via a Transformer-based architecture to capture long-range dependencies. Extensive experiments on multi-species datasets demonstrate that GLIMP achieves state-of-the-art performance, with an Accuracy of 96.3% and an F1-score of 96.3% (on A. thaliana), substantially improving recognition robustness and generalization across diverse promoter types.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy

