分类逻辑 通过神经网络进行两样样本测试,以区分近多重密度
Xiuyuan Cheng1, Alexander Cloninger2
1Department of Mathematics, Duke University, Durham, NC, 27708 USA.
概括
本研究引入了一种用于两个样本问题的新型网络逻辑测试,有效地区分数据密度. 该方法为大型数据集提供了计算优势,并且与现有技术相比显示出更高的性能.
科学领域:
- 机器学习 机器学习
- 统计推理 统计推理
- 数据科学数据科学数据科学
背景情况:
- 经典的两个样本问题旨在使用有限的数据样本来区分两个概率密度.
- 最近在生成对抗网络和变异学习方面的进展表明,分类网络在解决这个问题方面具有潜力.
- 现有的基于网络的方法为大型数据集提供了可扩展性,但需要进一步分析准确性和错误限制.
研究的目的:
- 开发和分析一个新的基于网络的统计数据,用于两个样本的问题,使用分类逻辑函数.
- 为了研究区分接近多重密度的逻辑函数的近似和估计错误.
- 建立对logit函数区分密度的能力的理论保证,考虑网络参数化和数据维度.
主要方法:
- 使用来自训练有素的神经网络的分类逻辑函数作为两个样本统计数据.
- 介绍神经网络对近多重积分近似的新理论结果,以分析逻辑函数错误.
- 证明 logit 函数能够在有足够的网络参数化和分析多重数据复杂性的情况下区分次指数密度.
主要成果:
- 网络logit测试显示了比以前基于网络的分类准确性测试更高的性能.
- 实验结果显示,与合成和现实世界数据集的内核最大平均差异测试进行了有利的比较.
- 理论分析证实了密度的可证明差异化和对多重数据的网络复杂性的降低.
结论:
- 拟议的网络逻辑测试是解决两个样本问题的有效和计算效率高的方法.
- 该方法可以很好地扩展到大型数据集,并比现有技术提供更好的性能.
- 理论见解为理解基于网络的密度差异化的能力和局限性提供了基础.
相关概念视频
McNemar's Test
278
McNemar's Test is a nonparametric statistical test used to determine if there is a significant difference in proportions between two related groups when the outcome is binary (e.g., yes/no, success/failure). It is beneficial when we have paired data, such as pre-test/post-test designs, where the same subjects are measured under two different conditions. The test is named after the statistician Quinn McNemar, who introduced it in 1947. It is commonly used in situations where subjects are...
278
The Anderson-Darling Test
742
The Anderson-Darling test is a statistical method used to determine whether a data sample is likely drawn from a specific theoretical distribution. Unlike parametric tests, it does not require assumptions about specific parameters of the distribution. Instead, it compares the sample's empirical cumulative distribution function (ECDF) with the cumulative distribution function (CDF) of the hypothesized distribution. Critical values for the test are specific to the chosen distribution rather...
742
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
142
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
142
Bonferroni Test
2.8K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.8K


