Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: Advancements in X-ray CT Tool Chain for Tree Core Analysis
Published on: September 22, 2023
Tree-based SIMCA for dealing with heterogeneous and sparse data
Robert van Vorstenbosch1, Frederik-Jan van Schooten2, Zlatan Mujagic3
1Department of Pharmacology and Toxicology, Maastricht University, the Netherlands; NUTRIM Institute of Nutrition and Translational Research in Metabolism, Maastricht University, the Netherlands; Institute of Risk Assessment Sciences, Population Health Sciences, Utrecht University, the Netherlands.
Background:
One Class Modelling (CM) is popular among chemometricians, but not well known among omics scientists in general. One issue is that typical CM approaches, including SIMCA, often result in unsatisfactory results due to e.g. large variation, centring and scaling issues, sparsity, outliers, and non-linearities in typical omics data. These effects can cause an inflated decision boundary (of the target class), thereby returning many false positives (of non-target cases). Tree-based techniques are by nature resistant to these challenges. In this study we explore tree-Based SIMCA variants in omics scenarios and compare to existing strategies.
Results:
We present a non-linear form of SIMCA by making use of sample proximities obtained through Unsupervised Random Forest and Isolation Forest (termed URF-SIMCA and IF-SIMCA). We compare accuracy of the algorithms with (traditional) SIMCA, one-class support vector machines, and isolation forest. This comparison was based on five (previously published) clinical omics datasets and the wine-dataset. URF-SIMCA showed superior behaviour. Using the pseudo-sampling principles, an interpretation could be made on the important features for the separation between the target and non-target classes. Using the wine-dataset, we empirically show that these directly relate to information obtained through two-class algorithms. Moreover, feature trajectories in the score- and orthogonal distance spaces further enable interpretability of the model.
Significance:
URF-SIMCA offers an easy to use extension of SIMCA, which deflates the variance of the target class, allowing for better separation. The increased modelling performance comes at the cost of feature interpretation, but this can be tackled using the pseudo-sampling principle.
More Related Videos
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
08:16Collecting and Processing Drone-based Remotely Sensed Data for Use in Forest Recovery Monitoring
Published on: October 24, 2025
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Statistical Methods for Analyzing Epidemiological Data
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...