基于机器学习的科学中的泄漏和可重复性危机
Sayash Kapoor1, Arvind Narayanan1
1Department of Computer Science and Center for Information Technology Policy, Princeton University, Princeton, NJ 08540, USA.
机器学习 (ML) 方法经常遭受数据泄露,导致过度乐观的结果. 纠正这些错误揭示了复杂的ML模型在许多科学应用中没有超过传统的逻辑回归 (LR).
科学领域:
- 量化科学 量化科学
- 计算社会科学 计算社会科学
背景情况:
- 机器学习 (ML) 方法在定量科学中越来越多地使用.
- 方法上的陷,如数据泄露,可能会损害基于ML的研究的可靠性.
- 在科学中应用ML时,可复制性问题是很重要的关注点.
研究的目的:
- 系统地调查基于机器学习的科学中的可重复性问题,重点关注数据泄露.
- 在机器学习研究中识别和分类不同类型的数据泄露.
- 提出一种测试和减轻数据泄露的方法.
主要方法:
- 在17个利用ML方法的科学领域进行系统的文献调查.
- 开发八种类型的数据泄露的详细分类.
- 介绍模型信息表,供研究人员进行泄漏测试.
- 在内战预测中,复杂的ML模型与传统的后勤回归 (LR) 进行了可复制性研究.
主要成果:
- 在17个领域的294篇论文中发现了数据泄露,往往导致了夸张的结论.
- 建立了八种数据泄露类型的综合分类.
- 模型信息表被提出作为评估泄漏的工具.
- 在一项内战预测案例研究中,经过校正的ML模型与LR模型相比,没有显著的性能优势.
结论:
- 数据泄露是基于机器学习的科学中普遍存在的问题,它破坏了可重现性,并导致性能指标膨胀.
- 拟议的分类学和模型信息表为识别和解决泄漏提供了一个框架.
- 复杂的ML模型在应用方法严格时,可能本质上不会超过更简单,更成熟的统计方法,如LR.
更多相关视频
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
相关概念视频
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Propagation of Uncertainty from Systematic Error
Random and Systematic Errors
Survival Tree
Building a Survival Tree
Constructing a...
Uncertainty in Measurement: Accuracy and Precision
Steps in Outbreak Investigation
