检测和缓解健康预测机器学习数据集转移的策略:系统审查
Gabriel Ferreira Dos Santos Silva1, Fabiano Novaes Barcellos Filho1, Roberta Moreira Wichmann2
1School of Public Health - University of São Paulo. Av. Dr. Arnaldo, 715 - Cerqueira César, São Paulo, SP 01246-904, Brazil.
识别和纠正机器学习中的数据集转移对于健康预测至关重要. 虽然方法存在,但没有一种方法是普遍适用的,这凸显了进一步的现实世界验证的需要.
科学领域:
- 在医疗保健中的机器学习
- 数据科学
- 生物医学信息学
背景情况:
- 数据集的转移对医疗保健中的机器学习 (ML) 模型的可靠性构成重大挑战.
- 确保ML预测的稳定性需要有效的策略来检测和减轻数据不一致性.
研究的目的:
- 在健康预测ML应用中全面审查识别和纠正数据集转移的方法.
- 综合有关医疗领域数据集转移检测和校正技术的当前文献.
主要方法:
- 在主要的科学数据库 (PubMed,IEEE Xplore,Scopus,Web of Science) 中进行系统的文献搜索.
- 包括2019年至2025年间发表的32项研究,重点关注ML,医疗保健和数据集转移.
- 基于数据集转移类型,检测/纠正策略,算法和性能影响的研究评估.
主要成果:
- 时间转移和概念漂移是最常见的数据集转移类型.
- 基于模型的监测和统计测试主导检测;再培训和特征工程是常见的校正方法.
- 现有的方法显示适度的解释性和可行性,但缺乏标准化指标和外部验证.
结论:
- 没有一个单一的数据集转移管理方法可以广泛地适用于所有健康ML应用.
- 目前这些技术在现实世界中的临床应用是有限的.
- 未来的研究应侧重于前性评估,子组分析和临床决策支持系统集成,以实现强大而公平的ML部署.
更多相关视频
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
相关概念视频
Steps in Outbreak Investigation
Regression Toward the Mean
Bias in Epidemiological Studies
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Statistical Methods for Analyzing Epidemiological Data
