超越alpha和omega:单测试可靠性估计器在单维连续数据中的准确性
1Kwangwoon University, Seoul, Korea. bene@kw.ac.kr.
Behavior research methods
|February 21, 2024
概括
像α系数这样的可靠性估计器是常见的,但这项研究发现标准化因子分析估计器通常是最准确的. 估计器的性能取决于数据特征,如样本大小和异常值.
科学领域:
- 心理测量 心理测量 心理测量
- 统计建模 统计建模
背景情况:
- 系数alpha是一个广泛使用的,但潜在的低于最佳的可靠性估计器.
- 因素分析 (FA) 估计器通常被推为优质的替代方案.
- 现有的文献表明,非标准化的估计器比标准化的估计器更准确.
研究的目的:
- 为了评估12种不同的可靠性估计器的准确性.
- 测试关于非标准化FA估计器优越性的常规智慧.
- 调查数据特征对估计器准确性的影响.
主要方法:
- 使用蒙特卡洛模拟来评估估计器的性能.
- 研究了12个不同的可靠性估计器.
- 数据特征如样本大小,项目数量和异常值被操纵.
主要成果:
- 几个估计器,包括FA和非FA类型,都超过了系数alpha.
- 一个标准化的FA估计器证明了最高的平均准确性.
- 标准化的估计器通常表现优于非标准化的估计器,这与之前的观点相反.
- 数据特征显著影响估计器的准确性;标准化的估计器在小样本和异常值方面表现出色,而非标准化的估计器在其他方面表现更好.
- 最大的下界估计器对3个项目准确,但对更多项目的可靠性高估.
结论:
- 没有单一的可靠性估计器在所有数据条件中普遍优越.
- 选择可靠性估计器应以特定数据特征为指导.
- 在某些条件下,标准化的FA估计器代表了一个非常准确的选择.
相关概念视频
Wilcoxon Rank-Sum Test
182
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
182
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Testing a Claim about Standard Deviation
2.5K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.5K
Friedman Two-way Analysis of Variance by Ranks
196
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
196
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


