在对联设计中,对二进制和多类F1 -scores的假设测试程序
Kanae Takahashi1, Kouji Yamamoto2, Aya Kuchiba3
1Department of Biostatistics, Hyogo Medical University, Hyogo, Japan.
Statistics in medicine
|August 1, 2023
概括
本研究开发了一种假设测试程序,用于在医学诊断中比较两个F1分数. 这种新方法使用对联研究设计和多变量中央极限定理来进行准确的统计推理.
科学领域:
- 医学诊断 医学诊断 医学诊断
- 统计方法 统计方法
- 机器学习评估 机器学习评估
背景情况:
- 医学测试对于诊断,查和风险预测至关重要.
- 像灵敏度,特异性和预测值这样的性能指标是二进制测试的标准.
- 准确度和回忆的和平均值F1得分越来越多地使用,用于多类问题的微和宏平均版本.
研究的目的:
- 开发一个假设测试程序来比较两个F1分数.
- 解决F1分数的统计推理方法上的差距,特别是在配对研究设计中.
- 为评估多类医学分类模型提供一个强大的统计框架.
主要方法:
- 开发一个假设测试程序来比较两个F1分数.
- 使用大样本多变量中心极限定理.
- 应用方法对配对研究设计进行医学试验性能评估.
主要成果:
- 开发了一种新的假设测试程序,用于比较对联研究中的F1分数.
- 该方法基于大样本多变量中心极限定理,确保统计有效性.
- 这有助于在医学中推进机器学习模型评估的统计推理.
结论:
- 开发的假设测试程序提供了一个统计学上合理的方法来比较F1分数在配对的医学研究设计.
- 这项研究增强了评估医疗诊断测试性能的统计工具包.
- 对F1分数的假设测试进行进一步的方法开发对于推进医学数据科学至关重要.
相关概念视频
Identifying Statistically Significant Differences: The F-Test
1.7K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
1.7K
Types of Hypothesis Testing
26.6K
There are three types of hypothesis tests: right-tailed, left-tailed, and two-tailed.
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...
26.6K
Bonferroni Test
2.8K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.8K
One-Way ANOVA: Unequal Sample Sizes
5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
Friedman Two-way Analysis of Variance by Ranks
240
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
240
Sign Test for Matched Pairs
162
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
To conduct the sign test, we first calculate the differences in...
162


