Related Experiment Video
Updated: Jul 1, 2026

14:56
Sample Preparation to Bioinformatics Analysis of DNA Methylation: Association Strategy for Obesity and Related Trait Studies
Published on: May 6, 2022
Model-based clustering of DNA methylation array data: a recursive-partitioning algorithm for high-dimensional data
E Andres Houseman1, Brock C Christensen, Ru-Fang Yeh
1Department of Biostatistics, Harvard School of Public Health, Boston, Massachusetts, 02115, USA. ahousema@hsph.harvard.edu
BMC Bioinformatics
|September 11, 2008
Summary
We developed a new algorithm for clustering DNA methylation array data. This method is reliable, computationally efficient, and accurately identifies methylation subgroups associated with tissue type and age.
Area of Science:
- Genomics
- Epigenetics
- Bioinformatics
Background:
- Epigenetics studies heritable gene function changes without altering DNA sequence.
- Cytosine methylation is a key epigenetic mechanism for gene silencing, often implicated in cancer.
- DNA methylation arrays, like Illumina GoldenGate, analyze thousands of cancer-related genes, but scalable clustering remains a challenge.
Purpose of the Study:
- To develop a novel, scalable, and reliable clustering algorithm for DNA methylation array data.
- To address the limitations of existing clustering methods for high-throughput epigenomic data.
Main Methods:
- A model-based recursive-partitioning algorithm was developed to navigate clusters within a beta mixture model.
- Simulations were used to compare the proposed method against nonparametric and conventional mixture model approaches.
Main Results:
- The proposed algorithm demonstrated superior reliability compared to nonparametric methods and comparable reliability to conventional mixture models.
- The method offers improved computational efficiency over traditional mixture model approaches.
- Application to normal tissue samples revealed clusters associated with tissue type and age.
Conclusions:
- The recursively-partitioned mixture model provides an effective and computationally efficient solution for clustering DNA methylation data.
- This approach facilitates the identification of biologically relevant subgroups within epigenomic datasets.
