Related Experiment Video
Updated: Jul 5, 2026

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Tutorial: Using Random Forest Analysis to Identify Auxiliary Variables of Missing Data
Stefany Coxe1,2, Amanda N Baraldi3, Timothy Hayes4
1Biostatistics Shared Resource, Cedars-Sinai Cancer, Los Angeles, CA, USA. stefany.coxe@csmc.edu.
None:
Missing data is a pervasive problem in research; prevention science is particularly vulnerable due to the study designs used and types of data collected. Recommended approaches to address missing data include full information maximum likelihood estimation and multiple imputation, both of which rely on identification of auxiliary variables related to missingness. The methodological literature recommends including as many potential auxiliary variables as possible, but in practice, that is often infeasible and a researcher must select a more limited number. Prior studies have shown that traditional methods for identifying auxiliary variables do not perform well when missingness follows a nonlinear functional form, but machine learning methods such as random forest analysis (RFA) perform well at successfully identifying correlates of missingness across a variety of missing at random patterns. RFA models can also provide measures of variable importance, allowing researchers to prioritize inclusion of the most relevant variables. Methods like RFA are less familiar to prevention researchers, which serves as an impediment for researchers to use RFA to identify auxiliary variables. This paper provides a tutorial to use RFA to identify correlates of missingness using a real data sample of = 215 participants with more than 100 variables measuring demographics, personality, and psychopathology. The tutorial demonstrates how to run an RFA (including using measures of variable importance to select the variables most related to missingness) and how to incorporate the selected auxiliary variables into an analysis.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Statistical Methods to Analyze Parametric Data: ANOVA
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares the...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...