绿洲:一个可解释的,有限样本的有效替代品皮尔森的X2科学发现
Tavor Z Baharav1,2, David Tse3, Julia Salzman4,5,6
1Eric and Wendy Schmidt Center, Broad Institute, Cambridge, MA 02142.
概括
一个新的统计测试,优化适应统计推断结构 (OASIS),为应急表提供高效和有效的分析. 在基因组数据分析方面,OASIS 卓越,能够检测出新型菌株,并在模拟中超越现有方法.
科学领域:
- 统计 统计 统计 统计
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 应急表在定量研究中至关重要,但现有的统计测试对于有限样本缺乏计算效率和统计有效性.
- 无参考基因组推断提出了挑战,需要强大的统计方法来分析计数数据.
研究的目的:
- 开发一个新的统计测试家族,用于意外表的优化适应统计推断结构 (OASIS).
- 提供适合有限观测的计算效率高和统计有效的测试,特别是在基因组应用中.
主要方法:
- OASIS在规范化数据矩阵中构建一个测试统计线性,利用闭式P值边界的度不等式.
- 导出OASIS测试统计数据的非对称分布,以验证有限样本边界.
- 将OASIS应用于基因组测序数据,以检测菌株和分析过度分散的数据.
主要成果:
- OASIS提供了应急表的可解释分解,有助于理解拒绝虚假假说.
- 实验证明了OASIS在基因组数据中的力量和可解释性,使得SARS-CoV-2和Mycobacterium结核病菌株的新检测成为可能.
- 模拟显示,OASIS对过度分散有强度,有效控制错误发现率,并在某些场景中优于皮尔森的千平方测试.
结论:
- OASIS代表了对应急表的统计测试的重大进步,提供了更高的效率和有效性.
- 该方法在基因组学中促进了新的应用,例如识别微生物菌株,这在以前是无法实现的.
- 与皮尔森的千平方测试等传统方法相比,OASIS表现出优越的性能和稳定性,特别是对于复杂的生物数据.
更多相关视频
20:24Characterization of Complex Systems Using the Design of Experiments Approach: Transient Protein Expression in Tobacco as a Case Study
Published on: January 31, 2014
16.5K
07:11Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
2.3K
相关概念视频
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Finding Critical Values for Chi-Square
2.9K
Consider a curve representing sample data drawn randomly from a normally distributed population. One must construct confidence intervals to estimate or to test a claim regarding the population standard deviation. For example, a 95% confidence interval covers 95% of the area under the curve, and the remaining 5% is equally distributed on either side of the curve. To achieve such confidence intervals, one must determine the critical values. The critical values are simply the values separating the...
2.9K
Fisher's Exact Test
488
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
488
Chi-square Analysis
38.2K
The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
38.2K
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
