Related Experiment Video
Updated: Dec 25, 2025

12:36
Chromatin Immunoprecipitation Assay for the Identification of Arabidopsis Protein-DNA Interactions In Vivo
Published on: January 14, 2016
21.0K
DDBJ Data Analysis Challenge: a machine learning competition to predict Arabidopsis chromatin feature annotations
Eli Kaminuma1, Yukino Baba2, Masahiro Mochizuki3
1Center for Information Biology, National Institute of Genetics.
Genes & Genetic Systems
|March 28, 2020
Summary
Machine learning competitions, like the DDBJ Data Analysis Challenge, effectively crowdsource models for complex life science tasks. This approach significantly improves predictive accuracy for DNA sequence annotation, benefiting experimental scientists.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Automating annotation analysis of large-scale next-generation sequencing data using machine learning (ML) is of high interest.
- Experimental life scientists often face challenges in finding ML collaborators.
- Crowdsourcing offers a potential solution to bridge this gap.
Framework:
- Investigated crowdsourced modeling for life science tasks via a machine learning competition: the DNA Data Bank of Japan (DDBJ) Data Analysis Challenge.
- Participants predicted chromatin feature annotations from DNA sequences using competing ML models.
- The challenge involved 38 participants and 360 model submissions.
Implementation:
- The top-performing model achieved an Area Under the Curve (AUC) score of 0.95.
- Overall model performance improved by an AUC of 0.30 during the competition.
- Top models incorporated external data, such as genomic location and gene annotations, demonstrating the value of domain knowledge.
Implications:
- Crowdsourced ML competitions can develop highly accurate models for experimental scientists lacking data science expertise.
- Incorporating domain knowledge into ML models led to significant performance improvements (5%-9% AUC).
- This approach democratizes access to advanced computational tools in life sciences research.
Related Concept Videos
Chromatin Immunoprecipitation- ChIP
12.0K
Chromatin immunoprecipitation, or ChIP, is an antibody-based technique used to identify sites on DNA that bind to transcription factors of interest or histone proteins. It also helps determine the type of histone modifications such as acetylation, phosphorylation, or methylation.
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
12.0K
Chromatin Position Affects Gene Expression
24.5K
Chromatin is the massive complex of DNA and proteins packaged inside the nucleus. The complexity of chromatin folding and how it is packaged inside the nucleus greatly influences access to genetic information. Generally, the nucleus' periphery is considered transcriptionally repressive, while the cell's interior is considered a transcriptionally active area.
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
24.5K
Genome Annotation and Assembly
20.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.4K

