Related Experiment Video
Updated: Aug 15, 2025

08:04
A Semiautomated ChIP-Seq Procedure for Large-scale Epigenetic Studies
Published on: August 13, 2020
3.5K
Comprehensive 100-bp resolution genome-wide epigenomic profiling data for the hg38 human reference genome
Ronnie Y Li1, Yanting Huang2, Zhiyue Zhao2
1Graduate program in Neuroscience, Emory University, United States.
Data in Brief
|December 30, 2022
Summary
This study offers high-resolution human epigenomic data from ENCODE, enabling machine learning for biological insights. The processed data and pipeline are publicly available for broad research applications.
Area of Science:
- Genomics
- Epigenetics
- Bioinformatics
Background:
- The human genome's epigenomic landscape is crucial for understanding gene regulation and biological processes.
- Existing epigenomic datasets often lack comprehensive coverage or standardized processing.
- The Encyclopedia of DNA Elements (ENCODE) consortium has generated vast amounts of sequencing data.
Purpose of the Study:
- To create a unified, high-resolution (100-bp) genome-wide epigenomic dataset for the human genome.
- To provide a standardized and accessible resource for researchers studying epigenetics and gene regulation.
- To develop a replicable data processing pipeline for epigenomic assays.
Main Methods:
- Processed raw sequencing data from five ENCODE assay types: DNase-seq, ChIP-seq (histone and TF), ATAC-seq, and Poly(A) RNA-seq.
- Filtered data for quality, merging technical replicates by averaging read counts into 100-bp bins across the genome.
- Developed a tabix-indexed dataset for efficient retrieval of read counts by genomic coordinates.
Main Results:
- Generated 6,305 high-quality, genome-wide epigenomic profiles at 100-bp resolution.
- Created a comprehensive dataset covering approximately 30 million 100-bp genomic bins.
- Provided an example of integrating epigenomic data with disease-risk variants from the GWAS Catalog.
Conclusions:
- The processed epigenomic data serve as valuable features for statistical and machine learning models.
- These data facilitate novel insights into gene expression, chromatin accessibility, and epigenetic modifications.
- The presented data processing pipeline is adaptable for other genome-wide assays, enhancing data accessibility and utility.
Keywords:
ATAC-seq, assay for transposase-accessible chromatin with sequencingBioinformaticsChIP-seq, chromatin immunoprecipitation followed by sequencingDNase-seq, DNase I hypersensitive site assay with sequencingENCODEENCODE, Encyclopedia of DNA ElementsEWAS, epigenome-wide association studyEpigenomicsGWAS, genome-wide association studyGenomicsHigh-throughput sequencingTF, transcription factorgnomAD, Genome Aggregation Database
