对稀疏PCA (研究) 的批判性评估:为什么 (应该承认) 重量不是负载
S Park1, E Ceulemans2, K Van Deun3
1Tilburg University, Methods and Statistics, Tilburg, The Netherlands. s.park_1@tilburguniversity.edu.
Behavior research methods
|August 4, 2023
概括
稀疏主要组件分析 (PCA) 方法可能会产生误导性的结果,如果数据生成和初始化策略不考虑稀疏的重量和稀疏的负载. 这项研究揭示了从常见实践中得出的过度乐观的结论.
科学领域:
- 多变量统计学 多变量统计学
- 数据挖掘 数据挖掘
- 机器学习 机器学习
背景情况:
- 主要组件分析 (PCA) 是减少维度和数据结构探索的基本技术.
- 标准PCA等同于权重,负载和正确的奇点向量,这一原则分解为稀疏的PCA变化.
- 稀少的PCA方法在重量或负载中引入零,从而导致经常被忽视的不同的数学解决方案.
研究的目的:
- 通过解决稀疏重量和稀疏载荷之间的忽视的区别,批判性地重新评估稀疏PCA方法.
- 调查数据生成方案和初始化策略对PCA性能稀疏的影响.
- 突出目前稀少的PCA应用中的潜在陷和可疑的研究实践.
主要方法:
- 通过整合专门为稀疏权重设计的数据生成方案,重新评估稀疏PCA.
- 评估各种初始化策略,超越标准的右单向量.
- 使用模拟和实证数据集对稀疏重量与稀疏负载方法进行比较分析.
主要成果:
- 在稀疏的PCA模拟中,常用的数据生成模型可能会导致过度乐观的性能结论.
- 在稀疏重量和稀疏装载方法之间的选择显著影响结果.
- 初始化策略对稀疏的PCA解决方案的融合和质量具有关键影响,可能导致局部最佳.
结论:
- 目前稀疏的PCA研究实践可能是有缺陷的,因为忽视了稀疏的重量和负载的独特特性.
- 强大的稀疏PCA需要仔细考虑数据生成机制和初始化技术.
- 经验证据表明,在稀疏的PCA中,方法选择带来了重大的实际后果.
相关概念视频
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Weighted Mean
5.2K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.2K
Calculating and Interpreting the Linear Correlation Coefficient
6.0K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.0K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Statistical Analysis: Overview
6.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.7K
Coefficient of Correlation
6.2K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.2K


