Related Experiment Video
Updated: May 20, 2026

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Multi-Level Variable Selection Using a BART-Enhanced Mixed-Effects Framework
Keming Zhang1, Yaoyao Li2, Jungang Zou3
1Department of Biostatistics, Brown University, Providence, Rhode Island, USA.
Abstract:
Selecting important individual- and cluster-level predictors has become increasingly critical in healthcare research, where data often exhibit hierarchical structures due to collection from multiple clusters. Mixed-effects models, which account for within-cluster correlation and between-cluster heterogeneity, are a natural approach for multilevel variable selection. However, currently available variable selection methods for multilevel data are predominantly based on mixed-effects models that impose restrictive parametric assumptions, potentially limiting their utility when the underlying relationships are nonlinear or involve interactions. While nonparametric methods have shown promise for variable selection in non-clustered data, they have been much less studied in the multilevel setting. Moreover, nonparametric methods that explicitly account for multilevel structure have largely been designed for prediction, rather than for simultaneous selection of relevant covariates at both the individual and cluster levels. To address these limitations, we propose a flexible, fully Bayesian unified framework for simultaneous variable selection of both fixed and random effects. Our framework integrates the nonparametric flexibility of Bayesian Additive Regression Trees (BART) for fixed-effect predictor selection with a hierarchical Bayesian component that identifies random-effect predictors via covariance decomposition and permutation strategies. To address scenarios common in multilevel data, where cluster-level covariates are constant within clusters and can induce near-collinearity and instability in selection, we further propose a computationally efficient two-step procedure. This method disentangles the contributions of individual- and cluster-level predictors, thereby mitigating collinearity and improving stability in variable selection. Comprehensive simulation studies demonstrate the effectiveness and robustness of our proposed methods across diverse scenarios. We further illustrate the practical utility of these approaches by applying them to a multilevel Alzheimer's disease dataset.
Related Concept Videos
Methods of Medium Optimization
Friedman Two-way Analysis of Variance by Ranks
Mechanistic Models: Compartment Models in Individual and Population Analysis
Experimental Designs
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Comparing the Survival Analysis of Two or More Groups

