通过Gwet的机会协议模型的最大概率来估计分级间的可靠性
Alek M Westover1, Tara M Westover2, M Brandon Westover2
1Massachusetts Institute of Technology, Boston, MA, USA.
概括
这项研究引入了最大概率卡帕 (maximum likelihood kappa),这是一个对评分器间可靠性 (IRR) 的公正估计器. 它纠正了现有的机会协议模型中的偏差,提高了IRR统计的准确性.
科学领域:
- 统计 统计 统计 统计
- 心理测量 心理测量 心理测量
- 数据科学数据科学数据科学
背景情况:
- 评价者间可靠性 (IRR) 量化了观察者之间的一致性.
- 科恩的卡帕是一个常见的IRR统计,但有局限性.
- 格韦特的协议统计提供了一个替代方案,但它有自己的偏见.
研究的目的:
- 为了解决现有的IRR统计中的局限性.
- 开发一个不偏见的估计机会协议.
- 为了介绍最大概率的卡帕 (kappa) 统计.
主要方法:
- 偶尔猜测模型的最大概率估计器的推导.
- 确定随机一致的概率与观察到的不一致率.
- 最大概率卡帕 (kappa) 统计的发展.
主要成果:
- 格韦特的机会协议公式在中间协议水平上有偏见.
- 最大概率卡帕 (Kappa) 提供了一个公正的IRR估计器.
- 随机一致的概率等于偶尔猜测模型下观察到的不一致率.
结论:
- 最大概率卡帕 (Kappa) 提供了一个理论上合理的IRR指标.
- 这个新的统计数据克服了科恩的卡帕和格韦特统计数据的局限性.
- 一个不偏见的IRR估计器对于可靠分析评级者判断至关重要.
相关概念视频
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
231
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
231
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Kendall's Coefficient of Concordance
173
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
173
Accuracy and Errors in Hypothesis Testing
156
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
156
Confidence Interval for Estimating Population Mean
7.2K
A point estimate of the population mean is obtained from a single sample. Such a point estimate does not represent a population well because it needs to account for variability in the population. Single point estimate can also be biased despite the sample being selected randomly. Thus, a point estimate is often unreliable. A confidence interval is needed to reduce this unreliability.
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
7.2K
Friedman Two-way Analysis of Variance by Ranks
99
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
99


