Related Experiment Video
Updated: Jun 6, 2026

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
Cross-study validation and guidance for resampling approaches for partial least squares discriminant analysis in
Aidan P Holman1, Dmitry Kurouski1
1Department of Biochemistry and Biophysics, Texas A&M University, College Station, TX, United States; Interdisciplinary Faculty of Toxicology, Texas A&M University, College Station, TX, United States.
Analytica Chimica Acta
|June 4, 2026
Summary
Resampling strategies significantly impact spectroscopic machine learning (SML) model selection and performance estimation. K-fold cross-validation offers the least biased performance estimates in SML.
Area of Science:
- Spectroscopic Machine Learning (SML)
- Chemometrics
- Statistical Learning
Background:
- Resampling strategies are crucial for model selection and performance evaluation in SML.
- Their influence is often underestimated compared to model choice.
Purpose of the Study:
- To systematically evaluate how common resampling methods affect model selection, predictive performance, and estimation bias.
- To compare resampling strategies within a fixed partial least squares-discriminant analysis (PLS-DA) framework.
Main Methods:
- Evaluation across twelve independent spectroscopic datasets.
- Model tuning using the one-standard-error (1SE) rule and highest Macro F1 (HMF1) selection.
- Performance assessment using repeated external validation and Bayesian Bradley-Terry analysis.
Main Results:
- Bootstrap and jackknife methods best preserved predictive performance.
- LOOCV, Venetian blinds, and some K-fold approaches balanced performance stability and model parsimony.
- Five-fold K-fold cross-validation yielded the least biased Macro F1 estimates, outperforming bootstrap and LOOCV.
Conclusions:
- Resampling strategy is a primary determinant of model selection and performance estimation accuracy in SML.
- K-fold cross-validation provides more reliable performance estimates than bootstrap or LOOCV.
- This study offers a framework for selecting optimal resampling methods to enhance SML model reliability and interpretability.
Related Concept Videos
Residuals and Least-Squares Property
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Sampling Plans
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Bootstrapping
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is small or...
Sampling Methods: Overview
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling.
In analytical chemistry, the choice of sampling...
In analytical chemistry, the choice of sampling...
Cluster Sampling Method
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
