在基于电子健康记录的机器学习预测模型中调整共变量错误分类
Shuang Yang1, Yonghui Wu1, Mei Liu1
1Department of Health Outcomes and Biomedical Informatics, University of Florida, Gainesville, Florida, USA.
这项研究引入了纠正电子健康记录数据错误的方法,改善了肺癌查的预测模型. 调整后的模型显示出比未调整的模型更好的准确性,减少了在没有手动审查的情况下的偏差.
科学领域:
- 生物医学信息学 生物医学信息学
- 医疗保健服务研究 医疗服务研究
- 医疗保健中的机器学习
背景情况:
- 电子健康记录 (EHR) 包含有价值的数据,但容易出现错误分类错误.
- 这些错误可以在预测模型中引入偏见,影响临床决策.
- 准确预测患者的结果,例如遵守肺癌查,至关重要.
研究的目的:
- 开发和评估用于调整EHR衍生的共变量中的错误分类错误的方法.
- 为了减少预测建模中的偏差,使用小组智能和个性化权重.
- 为了提高肺癌查依从性的预测模型的准确性.
主要方法:
- 开发了使用基于灵敏度和特异性的群智和个性化权重来调整EHR共变量的方法.
- 应用后勤回归,XGBoost和神经网络来预测肺癌查的坚持.
- 利用自然语言处理 (NLP) 来提取肺-RADS类别,并使用内核和多项回归权重进行调整.
- 将调整模型与未调整 (天真) 和真值 (预言) 模型进行比较,使用接收器操作特征 (AUROC) 曲线下的面积.
主要成果:
- 调整后的模型在各种验证集大小 (10%,20%,30%) 中显著优于原始模型.
- 与原始模型相比,AUROC的改进率在0.3%至10.4%之间.
- 与Oracle模型相比,调整后的模型将性能差距降低到2.0%-7.5%.
- 个体化权重显示出比群体权重更精确的错误纠正.
结论:
- 开发的框架有效地减轻了EHR衍生的共变量中的错误分类偏差.
- 这些方法提高了肺癌查遵守率的预测准确性,而不需要广泛的手动数据审查.
- 这种方法提供了一个可扩展的解决方案,用于提高使用真实世界临床数据的预测模型的可靠性.
更多相关视频
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
相关概念视频
Confounding in Epidemiological Studies
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Bias in Epidemiological Studies
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data
