Related Experiment Video
Updated: Mar 26, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
16.5K
A Regression Equation for the Parallel Analysis Criterion in Principal Components Analysis: Mean and 95th Percentile
Multivariate Behavioral Research
|January 23, 2016
Summary
This study introduces a new regression equation for parallel analysis, improving the accuracy of determining the number of principal components in factor analysis. The equation predicts mean eigenvalues and 95th percentile values from random data, aiding component retention decisions.
Area of Science:
- Statistics
- Psychometrics
- Data Analysis
Background:
- Parallel analysis is a common Monte Carlo method for determining the number of factors or components.
- Existing methods may have limitations in accuracy and reliance on chance.
Purpose of the Study:
- To develop a more accurate regression equation for predicting parallel analysis values.
- To improve the decision-making process for retaining principal components in factor analysis.
Main Methods:
- Developed a regression equation using random data sets (5-50 variables, 50-500 subjects).
- Predicted mean eigenvalues and 95th percentile eigenvalues from random data matrices with unities in diagonals.
- Validated the equation's accuracy against previous methods.
Main Results:
- The new regression equation accurately predicts parallel analysis values (multiple correlations ≥ .95).
- It is more accurate than a previously published equation for predicting mean eigenvalues.
- The equation also effectively predicts the 95th percentile of eigenvalue distributions.
Conclusions:
- The proposed regression equation offers a more reliable tool for parallel analysis in principal components analysis.
- This enhances the accuracy of determining the optimal number of components to retain.
- Easy-to-use tables of regression weights are provided for practical application.
Related Concept Videos
Residuals and Least-Squares Property
9.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.8K
Regression Toward the Mean
7.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.3K
Calculating and Interpreting the Linear Correlation Coefficient
8.5K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
8.5K
Percentile
9.7K
A percentile indicates the relative standing of a data value when data are sorted into numerical order from smallest to largest. It represents the percentages of data values that are less than or equal to the pth percentile. For example, 15% of data values are less than or equal to the 15th percentile.
9.7K
Outliers and Influential Points
6.7K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.7K
Residual Plots
6.7K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
6.7K

