错误发现率的确切积分公式和错误发现比例的方差
Rovshan G Sadygov1, Justin X Zhu1, Henock M Deberneh1
1Department of Biochemistry and Molecular Biology, The University of Texas Medical Branch, 301 University Blvd, Galveston, Texas 77555, United States.
Journal of proteome research
|May 29, 2024
概括
本研究提供了正假发现比例 (pFDR) 和多重假设测试中的差异的确切公式. 新的积分表达式提供了更好的准确性,特别是对于大规模数据分析中较小的假设集.
科学领域:
- 生物信息学和计算生物学
- 统计遗传学 统计遗传学
- 高通量欧米克数据分析数据分析
背景情况:
- 多重假设测试对于大规模的奥米克数据分析 (蛋白质组学,转录组学,代谢组学) 是至关重要的.
- 错误发现率 (FDR) 和积极的FDR (pFDR) 是标准的错误控制措施.
- 假发现比率 (FDP) 的预期 pFDR 经常被预期的比率所近似,但这种转换的条件尚不清楚.
研究的目的:
- 为期望 (pFDR) 和FDP的变量推导出精确的积分表达式.
- 调查pFDR常用的近似值的有效性和局限性.
- 为大规模的假设测试提供准确的错误估计方法.
主要方法:
- 对pFDR和FDP变异的精确积分表达式的导出.
- 对近似的分析 (预期比) 作为积分公式的局限性情况.
- 用特定数量的零假设计算pFDR的复杂度公式的开发.
- 在蛋白质组学中用于类鉴定的FDP差异的近似值.
主要成果:
- 成功地获得了pFDR和FDP变异的精确积分表达式.
- 广泛使用的预期近似比率是大样本大小的积分公式的具体情况.
- 模拟表明,对于少量的假设,积分表达式比近似更准确.
- 对于较大的样本大小,这两种方法的结果是可比的.
结论:
- 衍生的积分表达式为pFDR计算提供了更准确的框架,特别是当假设数量很小时.
- 这项工作阐明了精确的pFDR与其共同近似之间的关系,为实际应用提供了指导.
- 这些发现加强了大规模的奥米克研究中的错误控制,包括蛋白质组学数据分析.
相关概念视频
Testing a Claim about Population Proportion
3.3K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
3.3K
Identifying Statistically Significant Differences: The F-Test
1.6K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
1.6K
F Distribution
3.7K
The F distribution was named after Sir Ronald Fisher, an English statistician. The F statistic is a ratio (a fraction) with two sets of degrees of freedom; one for the numerator and one for the denominator. The F distribution is derived from the Student's t distribution. The values of the F distribution are squares of the corresponding values of the t distribution. One-Way ANOVA expands the t test for comparing more than two groups. The scope of that derivation is beyond the level of this...
3.7K
Testing a Claim about Standard Deviation
2.4K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.4K
Margin of Error
4.1K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
4.1K
Fisher's Exact Test
471
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
471


