Related Experiment Video
Updated: Jul 9, 2026

Targeted DNA Methylation Analysis by Next-generation Sequencing
Published on: February 24, 2015
Fast matrix completion in epigenetic methylation studies with informative covariates
Mélina Ribaud1, Aurélie Labbe1, Khaled Fouda1
1Department of Decision Science, HEC Montreal, 3000 chemin de la Cote Ste Catherine Montréal, QC H3T 2A7 Montreal, Canada.
Abstract:
DNA methylation is an important epigenetic mark that modulates gene expression through the inhibition of transcriptional proteins binding to DNA. As in many other omics experiments, the issue of missing values is an important one, and appropriate imputation techniques are important in avoiding an unnecessary sample size reduction as well as to optimally leverage the information collected. We consider the case where relatively few samples are processed via an expensive high-density whole genome bisulfite sequencing (WGBS) strategy and a larger number of samples is processed using more affordable low-density, array-based technologies. In such cases, one can impute the low-coverage (array-based) methylation data using the high-density information provided by the WGBS samples. In this paper, we propose an efficient Linear Model of Coregionalisation with informative Covariates (LMCC) to predict missing values based on observed values and covariates. Our model assumes that at each site, the methylation vector of all samples is linked to the set of fixed factors (covariates) and a set of latent factors. Furthermore, we exploit the functional nature of the data and the spatial correlation across sites by assuming some Gaussian processes on the fixed and latent coefficient vectors, respectively. Our simulations show that the use of covariates can significantly improve the accuracy of imputed values, especially in cases where missing data contain some relevant information about the explanatory variable. We also showed that our proposed model is particularly efficient when the number of columns is much greater than the number of rows-which is usually the case in methylation data analysis. Finally, we apply and compare our proposed method with alternative approaches on two real methylation datasets, showing how covariates such as cell type, tissue type or age can enhance the accuracy of imputed values.
More Related Videos
Related Concept Videos
Multiple Allele Traits
Mutation, Gene Flow, and Genetic Drift
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Multiple Allele Traits
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Genetic Variation
Genes exist in different versions called alleles, which...

