Related Experiment Video
Updated: Jun 25, 2026

Construction of Models for Nondestructive Prediction of Ingredient Contents in Blueberries by Near-infrared Spectroscopy Based on HPLC Measurements
Published on: June 28, 2016
Enhancing spectral interpretability and band selection prior to prediction model development via CAKE
Minh-Quan Nguyen1, Mizuki Tsuta2, Mito Kokawa3
1Graduate School of Science and Technology, University of Tsukuba, 1-1-1 Tennodai, Tsukuba, Ibaraki, 305-8572, Japan; Institute of Food Research, National Agriculture and Food Research Organization, 2-1-12 Kannondai, Tsukuba, Ibaraki, 305-8642, Japan.
Abstract:
A direct, singular causal relationship between the objective variable and spectral data underpins reliable and robust prediction models that avoid spurious correlations. However, machine learning models often lack causal interpretability due to their "black-box" nature. To address this, we developed the Causal Analysis via Kernel Estimation (CAKE), a framework that reveals single-component bands using only variables from the calibration model. CAKE employs an information-theoretic approach by calculating the differences in mutual information between variables and regression residuals to indicate causal direction. A Kernel Density Estimation (KDE) classifier further distinguishes causal structures prone to spurious correlations. The framework was optimized on simulated data and validated on real-world spectral measurements. Seven causal structures were constructed for the simulated data, while real-world data came from near-infrared and fluorescence spectroscopy of three-solvent mixtures (dimethyl sulfoxide, ethylene glycol, and glycerol). Three causality types - spectral bands with a single direct cause, hidden causes, and confounders - were distinctly characterized by density functions within the framework, enabling the accurate classification of all single-component bands into their respective causal structures. Compared with post-hoc approaches, such as external model validation and common variable selection indices, CAKE identifies reliable spectral bands that exclude spurious correlations across diverse spectroscopic applications while operating independently of the calibration process and requiring no prior knowledge of pure spectra. Although CAKE is designed as a pre-calibration method rather than a prediction-optimized variable selector, it achieves improved predictive accuracy, demonstrating that causal reliability and practical performance are not mutually exclusive.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Bandpass Sampling
A bandpass signal has a spectrum with a lower frequency limit, denoted as ω1, and an upper frequency limit, denoted as ω2. The spectrum...