在对eQTL数据的分组假设测试中控制错误发现
Pratyaydipta Rudra1, Yi-Hui Zhou2, Andrew Nobel3
1Department of Statistics, Oklahoma State University, Stillwater, OK, USA. prudra@okstate.edu.
BMC bioinformatics
|April 11, 2024
概括
我们开发了Z-REG-FDR,这是一种快速的方法,用于控制表达量化特征位置 (eQTL) 分析中分组假设的错误发现率 (FDR). 这种方法提供了更好的统计能力,适合大规模的基因组研究.
科学领域:
- 统计基因组学 统计基因组学
- 遗传流行病学遗传流行病学
- 生物信息学是一种生物信息学.
背景情况:
- 表达量的特征位点 (eQTL) 分析识别了影响基因表达的遗传变异.
- 基因水平的eQTL测试是一种对生物学理解至关重要的分组假设策略.
- 现有的控制组测试错误率的方法可能缺乏功率或适用于eQTL数据.
研究的目的:
- 开发一种新的方法来控制eQTL分析的分组假设测试中的错误发现率 (FDR).
- 使用随机效应组件来解决eQTL数据中效应大小的异质性.
- 为现有的FDR控制方法提供一个计算效率高的替代方案.
主要方法:
- 提出了随机效应模型和测试程序,用于在经验贝叶斯框架下进行集团级FDR控制 (REG-FDR).
- 引入了Z-REG-FDR,这是REG-FDR的近似,仅使用Z统计数据来提高计算效率.
- 使用模拟和真实eQTL数据评估方法性能.
主要成果:
- 在eQTL分析中,REG-FDR和Z-REG-FDR有效地控制了eQTL分析中的组合假设的FDR.
- Z-REG-FDR显示了与REG-FDR可比的统计性能.
- 与REG-FDR相比,Z-REG-FDR提供了显著提高的计算速度.
结论:
- 对于eQTL分析,Z-REG-FDR提供了统计能力和FDR控制的有利平衡.
- 该方法的速度和使用总结统计数据的能力使其对大规模统计基因组学非常实用.
- 在eQTL研究和相关的基因组研究中,Z-REG-FDR是对分组假设测试的宝贵工具.
相关概念视频
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Accuracy and Errors in Hypothesis Testing
198
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
198
Statistical Hypothesis Testing
1.9K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.9K
Decision Making: Traditional Method
4.0K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.0K
Errors In Hypothesis Tests
4.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.2K
Significance Testing: Overview
3.4K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
3.4K


