Related Experiment Video
Updated: May 25, 2026

O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
[A novel approach to NIR spectral quantitative analysis: semi-supervised least-squares support vector regression
1College of Information and Electrical Engineering, China Agricultural University, Beijing 100193, China. lilincau@gmail.com
A new semi-supervised LS-SVR model effectively analyzes near-infrared spectral data, utilizing samples with and without chemical values for improved quantitative analysis accuracy in tobacco composition.
Area of Science:
- Near-infrared spectroscopy
- Chemometrics
- Machine learning
Context:
- Quantitative analysis of chemical compositions using near-infrared (NIR) spectral data is limited by the availability of samples with accurately measured chemical values.
- Traditional mathematical models often exclude samples lacking chemical data, potentially reducing model performance.
- Developing robust models that leverage all available data, including unlabeled samples, is crucial for advancing NIR spectral quantitative analysis.
Purpose:
- To propose and evaluate a novel semi-supervised least squares support vector regression (S2 LS-SVR) model for near-infrared spectral quantitative analysis.
- To enhance the utilization of both labeled (with chemical values) and unlabeled (without chemical values) samples in chemometric modeling.
- To improve the accuracy and efficiency of predicting chemical compositions in complex samples.
Summary:
- A semi-supervised LS-SVR (S2 LS-SVR) model was developed, building upon the LS-SVR framework to incorporate samples lacking chemical values.
- The S2 LS-SVR model training is computationally efficient, equivalent to solving a linear system.
- Models were constructed for predicting total sugar, reducing sugar, total nitrogen, and nicotine content in flue-cured tobacco using PLS regression, LS-SVR, and the proposed S2 LS-SVR.
Impact:
- The S2 LS-SVR model demonstrated superior performance compared to PLS regression and standard LS-SVR, achieving average relative errors as low as 6.11% and high correlation coefficients (up to 0.9741) for tobacco composition analysis.
- This approach significantly improves the feasibility and efficiency of quantitative analysis in scenarios with limited labeled data.
- The findings highlight the potential of semi-supervised learning in chemometrics for more accurate and comprehensive spectral data analysis.
Related Concept Videos
Quantitative Analysis
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the method...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Applications of IR Spectroscopy: Overview
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
