Related Experiment Video
Updated: Jun 28, 2026

11:24
Direct Protein Delivery to Mammalian Cells Using Cell-permeable Cys2-His2 Zinc-finger Domains
Published on: March 25, 2015
Predicting DNA recognition by Cys2His2 zinc finger proteins
Anton V Persikov1, Robert Osada, Mona Singh
1Lewis-Sigler Institute for Integrative Genomics and Department of Computer Science, Princeton University, Princeton, NJ 08544, USA.
Bioinformatics (Oxford, England)
|November 15, 2008
Summary
We developed a new computational method using support vector machines (SVMs) to predict Cys(2)His(2) zinc finger (ZF) protein DNA binding. This approach improves accuracy by including data on weak and non-binding interactions.
Area of Science:
- Computational biology
- Bioinformatics
- Molecular biology
Background:
- Cys(2)His(2) zinc finger (ZF) proteins are the largest class of eukaryotic transcription factors.
- Their modular structure and conserved protein-DNA interface enable computational prediction of DNA-binding preferences.
- The canonical model for ZF protein-DNA interaction involves four amino acid nucleotide contacts per ZF domain.
Purpose of the Study:
- To develop and evaluate a novel computational approach for predicting Cys(2)His(2) zinc finger (ZF) protein DNA binding preferences.
- To investigate the utility of support vector machines (SVMs) incorporating non-binding and relative binding information.
Main Methods:
- Constructed a high-quality, literature-derived experimental database of ZF-DNA binding examples.
- Utilized support vector machines (SVMs) with linear and polynomial kernels to predict ZF protein-DNA binding.
- Incorporated information on weak and non-binding protein-DNA pairs, and relative binding affinities.
Main Results:
- The polynomial SVM demonstrated superior performance compared to previous prediction methods and the linear SVM.
- The results suggest potential dependencies between contacts in the canonical binding model.
- The approach incorporating non-binding and relative binding data shows significant potential for predicting protein-DNA interactions.
Conclusions:
- Support vector machines, particularly with polynomial kernels, are effective for predicting ZF protein-DNA binding.
- Including data on non-binding and relative binding interactions enhances prediction accuracy.
- Further refinement of the underlying structural model may lead to improved prediction performance.
Related Concept Videos
Single-Strand DNA Binding Proteins
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Cis-regulatory Sequences
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...

