文本挖掘以提取健康结果的适度性能可能几乎足够用于高质量的预后模型开发
Zwierd Grotenhuis1, Pablo J Mosteiro2, Artuur M Leeuwenberg3
1Department of Information and Computing Sciences, Utrecht University, The Netherlands; Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, The Netherlands.
使用文本挖掘来提取预后模型的患者结果显示出希望,即使质量较差的数据. 然而,风险预测可能需要重新校准以获得准确性.
科学领域:
- 医疗信息学 医疗信息学
- 临床流行病学临床流行病学
- 数据科学数据科学数据科学
背景情况:
- 预后模型估计患者对未来健康结果的风险.
- 训练这些模型需要具有特征和结果的历史患者数据.
- 结果数据通常是从临床笔记中提取的,使用文本挖掘,当它不结构化时.
研究的目的:
- 调查文本挖掘质量对预测模型性能的影响.
- 评估文本挖掘结果数据对开发预测模型的有用性.
主要方法:
- 进行了一项模拟研究,使用了医院死亡率预测的案例研究.
- 使用不同质量的文本挖掘模型提取的结果数据开发和评估了预测模型.
主要成果:
- 一个低质量的文本挖掘模型 (F1得分 ≈0.50) 能够训练出具有良好的区分能力的预测模型 (AUC ≈0.80).
- 预后模型的风险校准在大多数设置中仍然不可靠,即使是高质量的文本挖掘 (F1 ≈ 0.80).
结论:
- 从不完美的文本挖掘模型中使用文本挖掘结果开发预测模型是可行的.
- 以这种方式开发的预测模型可能需要使用手动提取的结果数据进行重新校准,以获得可靠的风险估计.
更多相关视频
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Cancer Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...
Statistical Methods for Analyzing Epidemiological Data
Kaplan-Meier Approach
