Related Experiment Video
Updated: Apr 21, 2026

O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
Assessment of the statistical significance of classifications in infrared spectroscopy based diagnostic models
David Pérez-Guaita1, Julia Kuligowski, Salvador Garrigues
1Centre for Biospectroscopy, School of Chemistry, Monash University, Clayton, 3800 VIC, Australia. bayden.wood@monash.edu.
Abstract:
Fourier transform infrared (IR) spectroscopy in combination with multivariate data analysis is a versatile tool that can be applied to disease diagnosis. However, a rigorous validation of the obtained models is necessary in order to obtain robust results. This work evaluates the advantages of the use of permutation testing for determining the statistical significance of the misclassification errors obtained from IR based diagnostic models through cross validation (CV). The model performance, estimated by CV, is compared to a distribution of CV-performance values obtained using randomly permuted class labels. The distribution of 'random CV-values' is considered as a null distribution and used to establish the significance of the model estimators obtained using real class labels. ATR-FTIR spectra of serum samples were classified using random forest (RF) classifiers according to two criteria, the tag number (a randomly assigned pseudo class membership) and the level of urea (real class). CV errors obtained were compared to the null distribution of CV errors from a permutation test and an independent validation set. The procedure was evaluated testing typical conditions leading to overoptimistic estimations provided by the CV like e.g. the size of subsamples used during CV, variable selection and the use of replicates. Results show that for the tag number (pseudo class), CV indicated classification errors between 23 and 33% depending on the subsample size employed. Those values were even lower when variable selection or replicates were used. However, permutation testing indicated that those CV errors were non-significant. In contrast, for sample classification according to their levels of urea, all cross validation errors were found to be significant. Although the proposed method is computationally intensive, it provides a simple way of calculating an empirical p-value of the CV-estimator, thus establishing the statistical significance and providing a feasibility indicator especially useful for studies where the number of samples is limited.
More Related Videos
11:05High-definition Fourier Transform Infrared FT-IR Spectroscopic Imaging of Human Tissue Sections towards Improving Pathology
Published on: January 21, 2015
10:25Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
Related Concept Videos
Infrared (IR) Spectroscopy: Overview
Different compounds display unique properties due to their...
IR Frequency Region: Fingerprint Region
Applications of IR Spectroscopy: Overview
IR Spectrum
Transmittance is defined as the ratio of the radiant power passing through a sample to that from the radiation's source. Multiplying the transmittance by 100 gives the percent transmittance (%T), which varies between 100% (no absorption) and 0%...
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
IR Spectroscopy: Molecular Vibration Overview
Stretching vibrations are vibrational motions that occur along the bond line, changing the bond length or distance between two bonded atoms. They are further distinguished as symmetric or asymmetric. In symmetric stretching, the...