检测模型不适合在结构方程建模与机器学习-一个概念验证
Melanie Viola Partsch1, David Goretzko1,2
1Department of Methodology and Statistics, University of Utrecht, Utrecht, The Netherlands.
Multivariate behavioral research
|November 3, 2025
概括
评估结构方程模型的合适性是一项挑战. 一种新的机器学习 (ML) 方法显示出准确评估多因素测量模型合适性的承诺,优于传统方法.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
背景情况:
- 结构方程建模 (SEM) 在心理学研究中被广泛使用.
- 使用固定的指数截止值来评估SEM合适性是有问题的,因为麻烦参数.
- 研究人员经常依赖这些切断,冒着错误的模型接受或拒绝的风险.
研究的目的:
- 开发一种基于机器学习 (ML) 的方法来评估多因素测量模型的合适性.
- 创建一种广泛适用的方法,尽量减少对干扰参数的依赖.
主要方法:
- 训练了一种ML模型,使用来自1,323,866个模拟数据集和确认因素分析模型的173个特征.
- 在1,659,386个独立测试观察结果上评估了ML模型的性能.
主要成果:
- 机器学习模型在检测各种条件下的模型 (错误) 匹配方面表现出高准确度.
- ML方法的表现优于传统的固定适合指数截止值.
- 轻微的错误规范,如单个残余相关性,对ML模型来说是具有挑战性的.
结论:
- 机器学习为改善SEM模型适合性评估提供了一个有希望的途径.
- 与传统技术相比,开发的ML方法显示出更高的性能.
- 需要进一步的研究来解决细微的模型错误规范的检测.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
277
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
277
Goodness-of-Fit Test
8.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
8.1K
Expected Frequencies in Goodness-of-Fit Tests
7.1K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis
237
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
237
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K


