在线虚假发现率控制LORD++和SAFFRON在阳性,局部依赖下
1Foundation Medicine Inc., Cambridge, Massachusetts, USA.
Biometrical journal. Biometrische Zeitschrift
|December 16, 2023
概括
新的研究表明,SAFFRON和LORD++方法可以在线控制错误发现率 (FDR),即使有依赖测试统计数据. 这扩大了它们在特定条件下超出修改FDR (mFDR) 控制范围的适用性.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 在线测试程序顺序分析假设,根据先前的测试统计数据调整显著性值.
- 像阿尔法投资,LORD++和SAFFRON这样的流行的方法在依赖 (CS) 条件下控制了修改错误发现率 (mFDR).
- 现有的文献主要显示LORD++和SAFFRON仅在独立性假设下控制传统的错误发现率 (FDR).
研究的目的:
- 证明SAFFRON和LORD++在局部非负的依赖条件下提供在线FDR控制.
- 将FDR控制保证扩展到在线测试中的自适应性停车规则.
- 正式描述条件超均假设对p值依赖性所施加的限制.
主要方法:
- 在线测试程序的理论分析.
- 将现有框架 (SAFFRON) 扩展到更广泛的依赖条件.
- 条件超均性假设的正式表征.
主要成果:
- 在本地非负依赖下,SAFFRON和LORD++确保在线FDR控制.
- 通过适应性停止规则保持FDR控制,包括在一定的拒绝次数后停止.
- 阿尔法投资被证明是SAFFRON的一个特殊案例,继承了这些FDR控制属性.
- 条件超均性假设在限制p值依赖性的作用已被正式定义.
结论:
- 这项研究扩大了在线FDR控制方法的理论保证,如SAFFRON和LORD++.
- 这些方法比以前建立的对统计依赖性的更强大.
- 这些发现对各种科学领域的顺序假设测试和适应性数据分析有影响.
相关概念视频
Critical Region, Critical Values and Significance Level
11.9K
The critical region, critical value, and significance level are interdependent concepts crucial in hypothesis testing.
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the...
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the...
11.9K
Significance Testing: Overview
3.4K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
3.4K
Frequency-dependent Selection
22.0K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
22.0K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Compacting Factor test
163
The compacting factor test is a method used to assess the workability of concrete. It is especially suitable for concrete mixes containing aggregates up to one and a half inches in size. This test involves specialized equipment consisting of two truncated cone-shaped hoppers and a cylinder, all with polished interior surfaces to minimize friction.
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
The procedure begins by placing concrete into the upper hopper without any compaction. Once filled, the bottom door of this hopper is opened,...
163
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K


