对高维数据进行非对称分布无独立性测试
Zhanrui Cai1, Jing Lei2, Kathryn Roeder2
1Faculty of Business and Economics, The University of Hong Kong.
Journal of the American Statistical Association
|December 9, 2024
概括
我们在高维数据中引入了独立性测试的新框架. 这种方法利用机器学习分类器来检测稀疏的依赖关系,为复杂的数据集提供了强大的工具.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 独立性测试对于变量选择,图形模型和因果推理至关重要.
- 由于缺乏分布或结构假设,高维和稀疏的数据对传统的独立性测试构成重大挑战.
研究的目的:
- 提出一个适用于高维,复杂数据的独立性测试的通用和强有力的框架.
- 开发一个测试统计数据的普遍,固定的高斯式零分布,独立于数据分布.
主要方法:
- 通过安装分类器来区分联合和产品分销来测试独立性的新框架.
- 使用来自机器学习的高级分类算法.
- 采用样本分割和固定的排列策略,以确保固定的高斯式零分布.
主要成果:
- 拟议的测试证明了在广泛的模拟中对现有方法的优势.
- 该框架有效地处理高维和稀疏的数据,优于当前的方法.
- 对单细胞测序数据集的成功应用,用于测试测量类型之间的独立性.
结论:
- 新框架为独立性测试提供了一种强大而灵活的方法,特别是对于复杂的高维数据.
- 该方法利用机器学习的能力提高了其在现代数据分析中的适用性.
- 普遍的零分布简化了解释,扩大了应用范围.
更多相关视频
13:55Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
Published on: February 3, 2013
18.3K
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
6.9K
相关概念视频
Introduction to Test of Independence
2.2K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.2K
Hypothesis Test for Test of Independence
3.5K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.5K
Test for Homogeneity
1.9K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
1.9K
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
F Distribution
3.7K
The F distribution was named after Sir Ronald Fisher, an English statistician. The F statistic is a ratio (a fraction) with two sets of degrees of freedom; one for the numerator and one for the denominator. The F distribution is derived from the Student's t distribution. The values of the F distribution are squares of the corresponding values of the t distribution. One-Way ANOVA expands the t test for comparing more than two groups. The scope of that derivation is beyond the level of this...
3.7K
One-Way ANOVA: Unequal Sample Sizes
5.7K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.7K
