Related Experiment Video
Updated: May 9, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Cloud-based uniform ChIP-Seq processing tools for modENCODE and ENCODE
Quang M Trinh1, Fei-Yang Arthur Jen, Ziru Zhou
1Ontario Institute for Cancer Research, MaRS Centre, South Tower, 101 College Street, Suite 800, Toronto, ON, M5G 0A3, Canada.
The modENCODE project offers a vast encyclopedia of functional genomic elements for C. elegans and D. melanogaster. New cloud-based resources and Galaxy workflows simplify data analysis, enabling reproducible research and efficient knowledge extraction from large genomic datasets.
Area of Science:
- Genomics
- Bioinformatics
- Model Organism Research
Background:
- The Model Organism ENCyclopedia of DNA Elements (modENCODE) project provides extensive functional genomic data for C. elegans and D. melanogaster.
- The large volume of modENCODE data (nearly 10 terabytes) presents challenges for researchers seeking to extract meaningful biological insights.
- Reinterpreting or combining modENCODE data with other datasets requires significant time and logistical effort.
Purpose of the Study:
- To address the challenges of analyzing large modENCODE datasets.
- To provide accessible and standardized computational resources for modENCODE data analysis.
- To facilitate reproducible research by offering consistent analytical frameworks.
Main Methods:
- Release of uniform computing resources on cloud platforms (Amazon Cloud, Bionimbus Cloud) integrated with Galaxy.
- Development of specific Galaxy workflows for analyzing ChIP-seq data, adhering to established quality control (QC) and peak calling standards.
- Creation of pre-configured machine images for cloud environments, including Galaxy, modENCODE data, and all necessary software dependencies.
Main Results:
- Uniform computing resources and Galaxy workflows are now available for analyzing modENCODE data.
- Cloud-based machine images simplify the setup and execution of complex genomic analyses.
- Standardized QC and peak calling ensure consistency across different research groups.
Conclusions:
- The provided resources establish a framework for consistent and reproducible analyses of modENCODE data.
- Researchers can now focus more on biological interpretation and less on data management and infrastructure setup.
- Enhanced accessibility to modENCODE data promotes efficient knowledge discovery in model organism genomics.
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Sanger Sequencing
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Complementary DNA
Complementary DNA
Genomics

