metadeconfoundR: Covariate analysis of high-dimensional cross-sectional omics data
Till Birkner1,2,3,4, Chia-Yu Chen1,2,3, Morgan Essex1,2,3
1Max-Delbrück Center for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany.
Motivation:
Identifying disease biomarkers from large molecular datasets is complicated by correlated and confounded signals like comorbidities and treatment regimens, batch effects, and cohort biases. These effects bias statistical inference and clinical conclusions. Robust methodologies are fundamental for reliable biomarker discovery.
Results:
metadeconfoundR is an R package for conservative biomarker discovery in (multi-)omics case-control datasets. It has a scalable two-step confounder-aware statistical framework for retaining only associations with independent support. It identifies covariate-naive univariate associations between omics features and metadata, then re-evaluates these associations using parallel post-hoc nested linear model testing to account for potential confounders. Confounded associations are flagged if they fully reduce to at least one other variable. metadeconfoundR supports parallel computation for large-scale datasets, offers visualization and tools for interpreting results and secondary analyses. We benchmark metadeconfoundR against state-of-the-art methods for identifying biomarkers using simulated ground truth derived from microbiome data, and demonstrate its ability to disentangle confounding effects while preserving statistical power, offering particular advantage when multiple covariates are present. metadeconfoundR functions for any -omics data type with continuous or categorical metadata/covariates.
Availability:
metadeconfoundR is available on CRAN (https://cran.r-project.org/web/packages/metadeconfoundR/) and GitHub (https://github.com/TillBirkner/metadeconfoundR).
Related Concept Videos
Confounding in Epidemiological Studies
Friedman Two-way Analysis of Variance by Ranks
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Longitudinal Studies
Two-Way ANOVA
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
Bias in Epidemiological Studies
