Sharing DNA-binding information across structurally similar proteins enables accurate specificity determination
Joshua L Wetzel1,2, Mona Singh1,2
1The Lewis-Sigler Institute for Integrative Genomics, Princeton, NJ 08544, USA.
Nucleic Acids Research
|November 29, 2019
Summary
This study introduces a new computational framework to improve DNA-binding specificity inference. By analyzing structurally similar proteins together, accuracy is significantly enhanced, leading to better predictions.
Area of Science:
- Computational Biology
- Bioinformatics
- Molecular Biology
Background:
- Thousands of protein-DNA interactions are experimentally assayed.
- Current computational methods infer DNA-binding specificities independently for each protein.
- This approach overlooks valuable interaction information from structurally similar proteins.
Purpose of the Study:
- To develop a novel framework for inferring DNA-binding specificities.
- To leverage interaction data from groups of structurally similar proteins simultaneously.
- To improve the accuracy of DNA-binding specificity predictions.
Main Methods:
- Developed a framework considering protein-DNA interactions for groups of structurally similar proteins.
- Devised constrained optimization and label propagation algorithms.
- Balanced individual protein observations with dataset-wide consistency.
Main Results:
- Tested approaches on two large, independent Cys2His2 zinc finger protein-DNA interaction datasets.
- Jointly inferring specificities dramatically improved accuracy within each dataset.
- Increased agreement between datasets and with an external standard.
Conclusions:
- Sharing protein-DNA interaction information across structurally similar proteins enhances prediction accuracy.
- The proposed framework offers a powerful method for accurate DNA-binding specificity inference.
- This approach improves consistency and agreement in protein-DNA interaction data analysis.
Related Concept Videos
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Single-Strand DNA Binding Proteins
16.4K
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
16.4K
Protein Complexes with Interchangeable Parts
2.8K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.8K
Cooperative Binding of Transcription Regulators
7.1K
Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome. Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form...
7.1K
Conservation of Protein Domains Over Different Proteins
13.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.9K
Ligand Binding Sites
14.8K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.8K


