线性混合模型和潜增长曲线模型用于受异常值污染的群组比较研究
Fabio Mason1, Eva Cantoni2, Paolo Ghisletta1
1Faculty of Psychology and Educational Sciences, University of Geneva.
Psychological methods
|February 15, 2024
概括
对于线性混合模型 (LMM) 和潜增长模型 (LGM) 的强大的估计方法在两组比较中处理主体内和主体间的异常值时显著提高了统计能力和I型错误率.
科学领域:
- 心理学统计 心理学统计
- 统计建模 统计建模
- 强大的统计数据.
背景情况:
- 线性混合模型 (LMM) 和潜增长模型 (LGM) 常用于分析两组比较中的纵向数据.
- 现有的LMM/LGM与异常值的研究主要涉及估计属性,而不是推断准确性.
- 需要在异常条件下评估统计能力和I型错误率.
研究的目的:
- 为了比较LMM和LGM的经典和强大的估计方法的性能,在主题内进行两组比较.
- 在各种异常情景下评估统计能力,I型错误率,置信区间 (CI) 覆盖范围,长度和平均绝对错误 (MAE).
- 确定最可靠的处理数据的方法,这些数据被污染了主体内和/或主体间的异常值.
主要方法:
- 进行了蒙特卡洛模拟实验,比较经典和强大的LMM/LGM估计器.
- 评估了Wald类型和引导式置信区间 (CI) 用于统计推断.
- 在四个条件下模拟数据:没有污染,主体内异常值,主体间异常值,以及两者.
- 包括强大的估计器,如S对LMM.
主要成果:
- 在没有异常值的情况下,经典和稳健的方法表现相似,尽管稳健的信贷机构的标称覆盖率略低.
- 在存在主体内和主体间异常值的情况下,可靠的估计器,特别是S,超过了经典方法.
- 在强大的LMM估计器上,百分位CI与野生启动是优越的,特别是在主体间的异常值.
- 发现LMM的古典沃尔德型CIs与异常值具有高度误导性.
结论:
- 强大的估计技术对于在数据包含异常值时,在LMM/LGM中准确推断至关重要.
- 百分点CI与野生引导应用于强大的LMM估计器提供了一个可靠的解决方案来处理复杂的异常情况.
- 研究人员应该考虑可靠的方法,以避免与受污染的数据进行两组纵向比较的误导性结果.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
186
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
186
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Friedman Two-way Analysis of Variance by Ranks
196
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
196
What Are Outliers?
3.8K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.8K


