贝叶斯因子,HDI-ROPE和频率等效测试都可以进行逆向工程-几乎完全-从彼此:回复林德等人. (2021年) 在2021年) 在2021年) 在2021年
Harlan Campbell1, Paul Gustafson1
1Department of Statistics, Faculty of Science, University of British Columbia.
Psychological methods
|March 21, 2024
概括
研究人员不应该偏爱基于经验性表现的贝叶斯式或频率主义等价性测试. 当对I型错误进行校准时,这两种方法都显示了类似的II型错误率,这表明方法选择取决于底层原则.
科学领域:
- 心理学统计 心理学统计
- 统计建模 统计建模
- 研究方法研究方法论.
背景情况:
- 相当性测试对于确定观察到的效应是否实际上可以忽略是至关重要的.
- 林德等人之前的研究. (2021) 推贝叶斯因子间隔零程序,超过频率主义TOST和贝叶斯HDI-ROPE.
- 这项建议的基础并没有明确与受控错误率挂.
研究的目的:
- 重新评估林德等人进行的模拟研究. (2021) 在标准化I型错误率下.
- 为了比较频率主义者TOST,贝叶斯式HDI-ROPE和贝叶斯系数间隔零程序的II型错误率.
- 评估贝叶斯式和频率式等效测试方法之间的实证性能差异.
主要方法:
- 由Linde等人进行的模拟研究的复制. (2021年). 在2021年.
- 将频率主义者TOST,贝叶斯式HDI-ROPE和贝叶斯系数间隔的零程序校准到一个共同的最大I型错误率.
- 在校准程序中分析II型错误率.
主要成果:
- 当校准到相同的最大I型错误率时,贝叶斯因子,HDI-ROPE和频率等效测试显示出几乎相同的II型错误率.
- 测试的贝叶斯式和频率式等效测试方法之间的经验性性能差异在受控错误率下是可以忽略不计的.
- 该研究表明,在等价性测试中,贝叶斯式和频率主义方法之间的选择应该以理论原则为指导,而不是经验性表现.
结论:
- 仅基于经验发现的频率主义或贝叶斯式测试是值得怀疑的.
- 选择同等性测试方法应与研究人员的基础统计学哲学保持一致.
- 当适当校准时,这些程序在很大程度上是可互换的,基于模拟结果,对彼此的论点会减少.
相关概念视频
Statistical Hypothesis Testing
1.9K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.9K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
128
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
128
Errors In Hypothesis Tests
4.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.2K
Significance Testing: Overview
3.4K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
3.4K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


