Related Experiment Video
Updated: Feb 16, 2026

08:16
Collecting and Processing Drone-based Remotely Sensed Data for Use in Forest Recovery Monitoring
Published on: October 24, 2025
648
Improved high-dimensional prediction with Random Forests by the use of co-data
Dennis E Te Beest1, Steven W Mes2, Saskia M Wilting3
1Department of Epidemiology and Biostatistics, VU University Medical Center, Amsterdam, 1007 MB, The Netherlands.
BMC Bioinformatics
|December 29, 2017
Summary
Auxiliary co-data enhances Random Forest prediction in high-dimensional settings. This method improves predictive accuracy by incorporating external information, outperforming standard Random Forests.
Area of Science:
- Machine Learning
- Bioinformatics
- Genomics
Background:
- High-dimensional prediction is challenging due to numerous variables and limited sample sizes.
- Standard Random Forests struggle in these complex data environments.
Purpose of the Study:
- To improve Random Forest predictive performance using auxiliary 'co-data'.
- To introduce a novel method for incorporating co-data into Random Forests.
Main Methods:
- Developed a co-data moderated Random Forest (CoRF) algorithm.
- Modified variable sampling probabilities using co-data, inspired by empirical Bayes.
- Applied CoRF to predict lymph node metastasis and cervical (pre-)cancer.
Main Results:
- CoRF demonstrated improved predictive performance in both case studies.
- Incorporated external p-values, gene signatures, and DNA copy number correlations for metastasis prediction.
- Utilized CpG island location, probe targeting, and related study p-values for cancer prediction.
Conclusions:
- Auxiliary co-data can significantly enhance Random Forest predictive capabilities.
- The CoRF method offers a robust approach for leveraging external information in high-dimensional prediction.
Related Concept Videos
Predicting Molecular Geometry
46.2K
VSEPR Theory for Determination of Electron Pair Geometries
46.2K
Random Error
9.9K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
9.9K
Random Variables
17.9K
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
17.9K
Randomized Experiments
9.1K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
9.1K
Random and Systematic Errors
15.4K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
15.4K
Prediction Intervals
3.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.4K

