RIFLE:从低订单边际计算的归算和强有力的推断
Sina Baharlouei1, Kelechi Ogudu1, Sze-Chuan Suen1
1University of Southern California.
统计推断的新框架RIFLE可以处理缺少的数据而无需归算. 它为回归和分类提供了可靠的估计,在具有高缺失率和低样本大小的具有挑战性的数据集中优于归算方法.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 缺失的数据在现实数据集中普遍存在,阻碍了统计推断和数据分析.
- 现有的归算方法经常失败,缺失率高,样本大小小,影响下游模型性能.
研究的目的:
- 开发一个新的统计推理框架用于回归和分类,以适应缺少的数据而没有归算.
- 引入RIFLE (通过低级瞬间估计进行强大的影响),用于分布强大的建模.
主要方法:
- RIFLE估计了底层数据分布的低序时刻和置信区间.
- 该框架专门用于线性回归和正常差分分析,并提供了趋同和性能保证.
- RIFLE还可以适应数据归算.
主要成果:
- 与最先进的归算和推理方法 (MICE,Amelia,MissForest,KNN-imputer,MIDA,Mean Imputer) 相比,RIFLE显示出更高的性能.
- 在数据集中,缺失值的百分比很高和/或样本大小小小的数据集中,超出性能尤为明显.
- 数字实验验证了框架的有效性.
结论:
- 在缺少数据的情况下,RIFLE为统计推理提供了一个强大的替代方案.
- 该框架适用于传统归算方法失败的场景.
- RIFLE提供了一个公开可用的解决方案,用于改进数据分析.
更多相关视频
09:23Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
Published on: August 16, 2017
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
相关概念视频
Distributions to Estimate Population Parameter
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Propagation of Uncertainty from Random Error
Assumptions of Survival Analysis
