归算质量对缺失值数据集的机器学习分类器的影响.
Tolou Shadbahr1, Michael Roberts2,3, Jan Stanczuk4
1Research Program in Systems Oncology, Faculty of Medicine, University of Helsinki, Helsinki, Finland.
分类不完整的数据集需要仔细的数据归算. 不良的归算质量显著降低了分类器的性能,突出了为可靠的机器学习结果优先考虑准确的归算方法的必要性.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 统计 统计 统计 统计
背景情况:
- 在不完整的数据集中对样本进行分类是常见的机器学习挑战.
- 现实世界的数据集往往含有缺失的值,需要在分类之前进行归算.
- 目前的重点是优化归算后的分类器性能.
研究的目的:
- 评估归算方法和缺失率对下游分类器性能的影响.
- 将现有的归算质量评估方法与使用切片瓦瑟斯坦距离的新方法进行比较.
- 分析训练在归算数据上的模型的稳定性和可解释性.
主要方法:
- 利用了三个模拟和三个现实世界的临床数据集,具有不同的失踪模式.
- 采用差异分析 (ANOVA) 来量化失踪率,归算和分类器选择的影响.
- 引入并评估基于切片的瓦瑟斯坦距离的差异得分,用于归算质量评估.
主要成果:
- 分类器的性能对测试数据缺失的百分比非常敏感.
- 常见的归算质量测量通常会导致数据分布与原始数据差异很大.
- 新的差异得分在匹配基础数据分布方面表现出卓越的表现.
- 当训练在不良归算数据上时,分类器模型的解释性就会受到损害.
结论:
- 数据归算的质量对于下游分类任务至关重要.
- 不良的归算可能会对分类器性能和模型可解释性产生相当大的负面影响.
- 优先考虑强大的归算质量评估对于可靠的机器学习应用程序至关重要.
更多相关视频
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
相关概念视频
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Detection of Gross Error: The Q Test
Survival Tree
Building a Survival Tree
Constructing a...
Quantifying and Rejecting Outliers: The Grubbs Test
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Improving Translational Accuracy
