Related Experiment Video
Updated: Jun 27, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
Published on: January 26, 2024
ESM2-BiMamba: A Length-Adaptive Hybrid Framework for Efficient Concurrent Prediction of DNA-Binding Proteins and
Yun Zhou1,2, Hanlei Qiu1, Dong Liu1,2
1College of Computer and Information Engineering, Henan Normal University, Xinxiang453007, China.
Abstract:
Accurate identification of DNA-binding proteins (DBPs) and their DNA-binding residue sites (DBSs) is essential for understanding gene regulatory processes. Despite recent progress achieved by protein language models, current methods still face fundamental limitations, including the quadratic computational burden of Transformer architectures, inadequate modeling of long-range dependencies, and reduced generalization on long or low-homology protein sequences. To address these challenges, we propose ESM2-BiMamba, a length-adaptive hybrid architecture for efficient and scalable protein-DNA interaction prediction. The model preserves the first 29 Transformer layers of the pretrained 33-layer ESM2 and replaces its top four layers with bidirectional Mamba state-space modules, enabling linear-time context propagation while maintaining rich sequence semantics. A sequence-length-adaptive dynamic chunking mechanism further reduces redundant computation and stabilizes long-range dependency modeling. To mitigate the distribution shift between pretraining and downstream tasks, a lightweight adapter is incorporated to enhance representation alignment. In addition, a dual-task prediction head comprising a protein-level DBP classifier and a residue-level DBS predictor allows the model to jointly capture global functional patterns and fine-grained binding-site signals. Extensive experiments on multiple standard benchmark data sets demonstrate that ESM2-BiMamba achieves superior performance in both DNA-binding protein identification and DNA-binding residue prediction, with notable advantages in processing long sequences and generalizing to low-homology targets.
More Related Videos
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay (EMSA) and DNA-affinity Precipitation Assay (DAPA)
Published on: August 21, 2016
09:17Structure-Based Simulation and Sampling of Transcription Factor Protein Movements along DNA from Atomic-Scale Stepping to Coarse-Grained Diffusion
Published on: March 1, 2022
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Single-Strand DNA Binding Proteins