Related Experiment Video
Updated: Jun 23, 2026

08:04
DNA Sequence Recognition by DNA Primase Using High-Throughput Primase Profiling
Published on: October 8, 2019
DNA-Prot: identification of DNA binding proteins from protein sequence information using random forest
K Krishna Kumar1, Ganesan Pugalenthi, P N Suganthan
1Institute for Neuro- and Bioinformatics, University of Lubeck, Lubeck 23538, Germany.
Journal of Biomolecular Structure & Dynamics
|April 24, 2009
Summary
A new random forest method, DNA-Prot, accurately identifies DNA-binding proteins (DNABPs) from protein sequences. This bioinformatics tool offers a reliable approach for understanding protein function in crucial cellular processes.
Area of Science:
- Bioinformatics
- Molecular Biology
- Computational Biology
Background:
- DNA-binding proteins (DNABPs) are crucial for fundamental cellular processes including DNA replication, repair, and transcriptional regulation.
- Current methods for identifying DNABPs primarily rely on protein structure, with limited options available for sequence-based identification.
Purpose of the Study:
- To develop and evaluate a novel random forest-based method, termed DNA-Prot, for the accurate identification of DNA-binding proteins directly from their amino acid sequences.
- To compare the performance of DNA-Prot against existing methods like DNAbinder using various benchmark datasets.
Main Methods:
- A random forest classification algorithm was trained and tested using curated datasets of known DNA-binding and non DNA-binding proteins.
- Performance evaluation involved calculating accuracy on training, testing, and independent benchmark datasets.
- Comparative analysis was conducted against the DNAbinder method using sequence-derived features.
Main Results:
- DNA-Prot achieved high accuracy, reaching 80.31% on the training set and 84.37% on the test set.
- Benchmarking on a large independent dataset (823 DNABPs, 823 non-DNABPs) demonstrated DNA-Prot's superior performance (81.83% accuracy) compared to DNAbinder (61.42%-63.5%).
- DNA-Prot consistently outperformed DNAbinder on additional benchmark datasets, highlighting its effectiveness.
Conclusions:
- DNA-Prot is an efficient and accurate computational tool for identifying DNA-binding proteins based solely on protein sequence information.
- The developed method provides a valuable resource for researchers studying protein function and cellular mechanisms.
- The DNA-Prot software and datasets are publicly available for further research and application.
Related Concept Videos
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Single-Strand DNA Binding Proteins
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...

