Related Experiment Video
Updated: Jan 28, 2026

Methods of Soil Resampling to Monitor Changes in the Chemical Concentrations of Forest Soils
Published on: November 25, 2016
Surrogate minimal depth as an importance measure for variables in random forests
Stephan Seifert1, Sven Gundlach1, Silke Szymczak1
1Institute of Medical Informatics and Statistics, Kiel University, University Hospital Schleswig-Holstein, Kiel, ermany.
A new variable selection method, Surrogate Minimal Depth (SMD), improves the identification of causal variables in high-dimensional omics data. SMD offers enhanced insights into predictor-outcome relationships compared to existing methods.
Area of Science:
- Bioinformatics
- Machine Learning
- Genomics
Background:
- Random forest is effective for omics data analysis, including classification, regression, and variable selection.
- Interpreting variable importance in complex, high-dimensional omics data is challenging due to intricate relationships between predictors.
Purpose of the Study:
- To introduce a novel variable selection approach, Surrogate Minimal Depth (SMD).
- To enhance the interpretability of variable importance in high-dimensional omics data by considering variable relationships.
Main Methods:
- Developed Surrogate Minimal Depth (SMD) by integrating surrogate variables into the Minimal Depth (MD) concept.
- Applied SMD to simulated correlation patterns to assess its ability to reconstruct relationships and improve variable selection.
Main Results:
- SMD successfully reconstructed simulated correlation patterns, demonstrating its capability to account for variable interdependencies.
- SMD exhibited higher empirical power in identifying causal variables compared to state-of-the-art methods and MD, with comparable stability.
- The method provides improved insights into the complex interplay between predictor variables and outcomes.
Conclusions:
- Surrogate Minimal Depth (SMD) is a promising advancement for variable selection in high-dimensional omics data.
- SMD offers a more nuanced understanding of predictor variable relationships and their impact on outcomes.
- The approach facilitates more accurate identification of causal variables in complex datasets.
Related Concept Videos
Random Variables
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Random Error
Randomized Experiments
Simple randomization
Simple...
Variables Affecting Phosphorescence and Fluorescence
Random and Systematic Errors
Uniform Depth Channel Flow

