在诊断潜阶级模型中检测偏离条件独立性假设的偏差:一个模拟研究
Yasin Okkaoglu1, Nicky J Welton2, Hayley E Jones2
1Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, UK. yasin.okkaoglu@bristol.ac.uk.
BMC medical research methodology
|December 5, 2024
概括
通常使用的统计工具用于评估诊断准确性,而没有黄金标准测试往往会失败. 剩余相关图和千方位统计数据对于检测依赖测试的功率很低,导致偏差估计.
科学领域:
- 生物统计学 生物统计学
- 统计建模 统计建模
- 诊断准确性的研究研究.
背景情况:
- 隐性类型模型估计诊断准确性没有黄金标准.
- 假设测试的独立性可以在存在依赖关系时对结果产生偏见.
- 剩余相关图和千平方统计数据用于检查测试依赖性.
研究的目的:
- 为了评估残余相关图和千平方统计的性能,在潜在类模型中识别依赖诊断测试.
- 在各种模拟场景中评估这些工具的可靠性.
主要方法:
- 模拟了504个数据集组合,样本大小,流行率,共变率,灵敏度和特异性各不相同.
- 从一个具有四个测试和测试1和2之间的依赖性模型中生成每组1000个数据集.
- 在贝叶斯框架中安装了条件独立模型,并分析了偏见,覆盖范围和缺陷检测.
主要成果:
- 剩余相关图和千平方统计仅在~10%~12%的时间内正确识别了依赖测试对.
- 这些工具错误地标记了测试3和4在超过50%的模拟中是依赖的.
- 在64-74%的模型中检测到缺乏整体匹配,条件独立模型显示对相关测试的偏差估计.
结论:
- 剩余相关性图和基平方统计数据对于识别诊断测试之间的条件依赖是不可靠的.
- 这些方法在检测整体模型匹配问题方面功率较低.
- 如果不考虑测试依赖性,则会导致显著的参数估计偏差.
相关概念视频
Hypothesis Test for Test of Independence
3.5K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.5K
Introduction to Test of Independence
2.2K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.2K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Statistical Hypothesis Testing
1.9K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.9K
Errors In Hypothesis Tests
4.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.2K
Comparing the Survival Analysis of Two or More Groups
150
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
150


