Related Experiment Video
Updated: Jan 4, 2026

O-cresol Concentration Online Measurement Based On Near Infrared Spectroscopy Via Partial Least Square Regression
Published on: November 8, 2019
Determining the number of components in PLS regression on incomplete data set
Titin Agustin Nengsih1,2, Frédéric Bertrand1, Myriam Maumy-Bertrand1
1IRMA, CNRS UMR 7501, Université de Strasbourg, 67084 Strasbourg, Cedex, France.
Partial least squares regression (PLS) with the NIPALS algorithm handles incomplete data effectively. The Q2-leave-one-out criterion is more reliable for selecting components in PLS regression with missing data than AIC or BIC.
Area of Science:
- Statistics
- Machine Learning
- Data Science
Background:
- Partial least squares regression (PLS) is a multivariate statistical method widely used in applied research.
- PLS regression effectively analyzes relationships between an outcome and multiple components.
- The NIPALS algorithm can estimate parameters even with incomplete datasets, but component selection with missing data is debated.
Purpose of the Study:
- To investigate the performance of the NIPALS algorithm in Partial Least Squares (PLS) regression with varying proportions and types of missing data.
- To compare different criteria for selecting the number of components in PLS regression when dealing with incomplete datasets.
- To evaluate the effectiveness of imputation methods prior to component selection in PLS regression.
Main Methods:
- The NIPALS algorithm was used to fit PLS regression models to datasets with simulated missing data (5% to 50%).
- Component selection was assessed using Q2-leave-one-out, AIC, and BIC criteria.
- Three imputation methods were compared: multiple imputation by chained equations, k-nearest neighbour imputation, and singular value decomposition imputation.
Main Results:
- The NIPALS algorithm's behavior was studied under various missing data scenarios.
- Q2-leave-one-out component selection demonstrated more reliable results compared to AIC and BIC criteria when applied to both incomplete and imputed datasets.
- The performance of imputation methods varied, impacting the subsequent component selection.
Conclusions:
- Q2-leave-one-out is a more robust criterion for selecting the number of components in PLS regression with missing data.
- Imputation strategies can influence the choice of components, highlighting the importance of method selection.
- Further research is needed to optimize imputation and component selection strategies for PLS regression in the presence of missing data.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Survival Tree
Building a Survival Tree
Constructing a...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

