Related Experiment Video
Updated: Aug 6, 2026

16:24
Analyzing and Building Nucleic Acid Structures with 3DNA
Published on: April 26, 2013
Weight matrices for protein-DNA binding sites from a single co-crystal structure
Robert G Endres1, Ned S Wingreen
1NEC Laboratories America, Inc., Princeton, New Jersey 08540, USA. rendres@princeton.edu
Summary
This study introduces an atomistic method to predict transcription factor DNA-binding sites using protein-DNA crystal structures. The Wang-Landau Monte Carlo algorithm efficiently generates accurate weight matrices, even with limited data.
Area of Science:
- Structural Biology
- Computational Biology
- Molecular Biophysics
Background:
- Transcription factors regulate gene expression by binding to specific DNA sequences.
- Identifying these DNA-binding sites is crucial for understanding cellular processes.
- Current methods often rely on numerous known binding sites, which are not always available.
Purpose of the Study:
- To develop an atomistic method for inferring transcription factor DNA-binding sites.
- To construct accurate weight matrices for DNA-binding site prediction from limited data.
- To demonstrate the utility of the Wang-Landau Monte Carlo algorithm in this context.
Main Methods:
- Utilizing X-ray co-crystal structures of protein-DNA complexes as a starting point.
- Employing the Wang-Landau Monte Carlo algorithm to efficiently sample high-affinity binding sites.
- Incorporating bound water molecules at the protein-DNA interface for accurate site recovery.
- Comparing results with the dead-end elimination algorithm for low-complexity cases.
Main Results:
- The atomistic method successfully infers DNA-binding sites and constructs accurate weight matrices.
- The Wang-Landau Monte Carlo algorithm efficiently samples binding sites, analogous to bioinformatics approaches.
- Inclusion of bound water is critical for recovering crystal binding sites.
- The method shows potential applicability even without a native co-crystal structure.
Conclusions:
- An atomistic computational approach can accurately predict transcription factor DNA-binding sites and generate weight matrices.
- The Wang-Landau Monte Carlo algorithm offers an efficient method for sampling binding sites.
- This method provides a valuable tool for studying gene regulation, especially when experimental data is scarce.
More Related Videos
Related Concept Videos
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Single-Strand DNA Binding Proteins
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Globular and Fibrous Proteins
Many proteins can be classified into two distinct subtypes - globular or fibrous. These two types differ in their shapes and solubilities.
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein-Protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...

