Related Experiment Video
Updated: Apr 11, 2026

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
PRIMED: predicting DNA binding residues by leveraging pre-trained protein language models.
Luoshu Zhang1, Xin Li1, Ruocen Song1
1J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida, Gainesville, FL, United States.
We developed PRIMED, a machine learning tool that accurately predicts DNA-binding residues in proteins by integrating multiple protein language models. This method enhances understanding of gene regulation and disease mechanisms.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Protein-DNA interactions are crucial for fundamental biological processes like gene regulation and genome stability.
- Identifying DNA-binding residues (DBRs) is essential for protein engineering and therapeutic design, but experimental methods are limited by low throughput.
- Computational approaches using protein sequence and structure offer scalable alternatives for DBR prediction.
Purpose of the Study:
- To develop and validate PRIMED, a novel machine learning framework for accurate prediction of DNA-binding residues.
- To integrate diverse protein representations from multiple protein language models for enhanced predictive performance.
- To assess PRIMED's generalizability across various protein datasets with different DBR frequencies.
Main Methods:
- PRIMED utilizes a multilayer perceptron to process concatenated protein representations derived from three distinct protein language models: ESM-2, ESM-3, and ESM-C.
- The framework integrates biochemical and structural properties encoded within these language model representations.
- Predictions were performed on benchmark datasets including Test-46, Test-129, CLAPE-DB, and a newly curated Test-10 K dataset.
Main Results:
- PRIMED achieved high performance on benchmark datasets, with an AUC of 0.92 and MCC of 0.64 on Test-46, and an AUC of 0.93 and MCC of 0.45 on Test-129.
- The model demonstrated strong generalizability on the large-scale Test-10 K dataset, performing competitively against existing methods.
- Results indicate PRIMED's effectiveness in predicting DBRs across proteins with varying DBR percentages.
Conclusions:
- Integrating diverse protein language model representations significantly improves the accuracy and transferability of DNA-binding residue predictions.
- PRIMED offers a powerful computational tool for advancing research in gene regulation, genome stability, and disease mechanisms.
- The framework's performance highlights the potential of advanced machine learning techniques in molecular biology.
More Related Videos
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
DNA as a Genetic Template
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Organization
The primary structure of a protein is its amino acid sequence....

