扩展信托推断:朝着统计推断的自动化过程进行统计推断
Faming Liang1, Sehwan Kim2, Yan Sun3
1Department of Statistics, Purdue University, West Lafayette, IN 47907, USA.
概括
扩展信托推断 (EFI) 恢复了R.A.费舍尔推断参数不确定性的目标. 这种可扩展的大数据方法使用深度神经网络和先进的计算来进行强大的统计推理和半监督学习.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 费舍尔最初的目标是通过信托推断推断参数不确定性,这在很大程度上被驳回了.
- 对于参数不确定性的强大的统计推理方法的追求仍然是一个活跃的研究领域.
研究的目的:
- 开发一种新的统计推理方法,即扩展信托推理 (EFI).
- 为了实现可扩展的大数据推断模型参数不确定性.
- 为半监督学习提供一个创新的框架.
主要方法:
- 通过使用随机梯度马尔科夫链蒙特卡罗联合推算观测错误.
- 用稀疏的深度神经网络 (DNN) 估计反函数.
- 利用先进的统计计算技术来实现可扩展性.
主要成果:
- 通过一致的DNN估计器,EFI确保了观察不确定性的适当传播到模型参数.
- 在参数估计方面,EFI表现出更高的准确性,特别是在异常值数据方面.
- EFI消除了对假设测试中的理论参考分布的需求,自动化推理.
结论:
- 扩展信任推理 (EFI) 成功实现了以可扩展的方式推断参数不确定性的目标.
- 在参数估计和假设测试方面,EFI比频率主义和贝叶斯方法具有显著的优势.
- 欧洲教育基金会 (EFI) 提出了一个适用于半监督学习任务的新框架.
相关概念视频
Identifying Statistically Significant Differences: The F-Test
1.6K
The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
1.6K
F Distribution
3.6K
The F distribution was named after Sir Ronald Fisher, an English statistician. The F statistic is a ratio (a fraction) with two sets of degrees of freedom; one for the numerator and one for the denominator. The F distribution is derived from the Student's t distribution. The values of the F distribution are squares of the corresponding values of the t distribution. One-Way ANOVA expands the t test for comparing more than two groups. The scope of that derivation is beyond the level of this...
3.6K
What are Estimates?
4.9K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
4.9K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
113
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
113
Fisher's Exact Test
308
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
308
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K


