最佳非参数推理与两个尺度分布式最近邻居.
Emre Demirkaya1, Yingying Fan2, Lan Gao1,2
1University of Tennessee Knoxville.
Journal of the American Statistical Association
|May 8, 2024
概括
这项研究引入了两级分布式最近邻近 (TDNN) 估计器,这是一种用于非参数平均回归的新偏差减小方法. TDNN实现了最佳的收率和非对称的正常性,使得有效的统计推断.
科学领域:
- 统计 统计 统计 统计
- 非参数统计的统计.
- 机器学习 机器学习
背景情况:
- 权重最近邻居 (WNN) 是一个灵活的非参数工具,用于平均回归.
- 分布式最近邻居 (DNN) 使用包装来创建WNN估计器.
- 现有的DNN方法缺乏分布结果和最佳的融合率,以实现顺的功能.
研究的目的:
- 解决DNN估计器的局限性,特别是在更高阶平滑性情况下的偏差问题.
- 开发一个减少偏差的DNN估计器,以实现最佳的非参数收率.
- 为拟议的估计器建立理论属性和实际实施工具.
主要方法:
- 通过将两个DNN估计器与不同的亚抽样尺度相结合,引入了一个偏差减少方法,创建了两级DNN (TDNN) 估计器.
- 提供了TDNN作为WNN估计器的等效表示,具有明确的,潜在的负面权重.
- 为DNN和TDNN估计器建立了非对称的正常性.
主要成果:
- 在第四阶顺条件下,TDNN估计器实现了最佳的非参数收率,克服了DNN的偏差限制.
- 理论分析证实了DNN和TDNN估计者的非对称正常性.
- 开发了TDNN的差异和分布估计器,使用jackknife和bootstrap技术进行实际推断.
结论:
- TDNN估计器比标准DNN显著改进,实现最佳的融合率,并使有效的统计推断成为可能.
- 理论结果和实际实施工具 (变异/分布估计器) 支持使用TDNN进行非参数回归.
- 该研究通过模拟和真实数据应用来证明TDNN的有效性.
相关概念视频
Introduction to Nonparametric Statistics
709
Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
One of...
709
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
124
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
124
Ranks
236
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
236
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Choosing Between z and t Distribution
2.8K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
2.8K
Friedman Two-way Analysis of Variance by Ranks
186
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
186


