Recognition models to predict DNA-binding specificities of homeodomain proteins

Ryan G Christensen1, Metewo Selase Enuameh, Marcus B Noyes

  • 1Department of Genetics, Washington University School of Medicine, St. Louis, MO 63108, USA.

Summary

Developing accurate protein-DNA interaction models is crucial for understanding gene regulation. This study introduces improved machine learning methods for predicting transcription factor binding specificities, outperforming existing approaches.

Related Concept Videos

Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Single-Strand DNA Binding Proteins01:03

Single-Strand DNA Binding Proteins

For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
Cooperative Binding of Transcription Regulators02:13

Cooperative Binding of Transcription Regulators

Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome.  Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form dimers that...
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
DNA as a Genetic Template02:05

DNA as a Genetic Template

Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
Histone Modification02:32

Histone Modification

The histone proteins have a flexible N-terminal tail extending out from the nucleosome. These histone tails are often subjected to post-translational modifications such as acetylation, methylation, phosphorylation, and ubiquitination. Particular combinations of these modifications form “histone codes” that influence the chromatin folding and tissue-specific gene expression.
Acetylation
The enzyme histone acetyltransferase adds acetyl group to the histones. Another enzyme, histone deacetylase,...