通过自然语言处理在基于机器学习的COVID-19死亡预测中的非结构化数据的增量值:一项比较研究
Rildo Pinto da Silva1, Antonio Pazin-Filho2
1Departamento de Clínica Médica, Faculdade de Medicina de Ribeirão Preto, Universidade de São Paulo, Ribeirão Preto, São Paulo, Brazil. rildo.silva@alumni.usp.br.
BMC medical informatics and decision making
|September 27, 2025
概括
将非结构化数据添加到机器学习模型中并没有显著改善COVID-19死亡率预测. 人类监督对于验证自然语言处理输出和选择这些模型的特性至关重要.
科学领域:
- 医疗信息学 医疗信息学
- 人工智能在医学中的应用
- 临床预测模型临床预测模型
背景情况:
- 机器学习模型越来越多地用于临床预测.
- 讨论的是非结构化临床数据对增强这些模型的价值.
- 很少有研究严格评估了非结构化数据对模型性能的影响.
研究的目的:
- 通过包括非结构化数据来评估用于医院死亡率预测的机器学习模型的性能改进.
- 将仅使用结构化数据的模型与包含非结构化数据的混合模型进行比较.
- 量化非结构化数据对预测准确性的影响.
主要方法:
- 一项回顾性研究比较了使用结构化数据的机器学习模型与添加非结构化数据的混合模型.
- 模型是为在第三级急诊医院诊断出COVID-19的患者开发的.
- 使用诸如接收器运行特征曲线下的面积 (AUC ROC),灵敏度和特异性等指标来评估性能.
主要成果:
- 最好的模型,一个随机森林,在非结构化数据中实现了0.9260的AUC ROC,比仅使用结构化数据的0.9170略有增加.
- 灵敏度从0.8108提高到0.8378,而特异性保持在0.8667.7的稳定水平.
- 与仅使用结构化数据的模型相比,这些性能增长在统计学上并不显著.
结论:
- 包括非结构化数据并没有显著提高机器学习模型对COVID-19死亡率的预测能力.
- 人类参与对于验证自然语言处理输出和选择相关的非结构化特征至关重要.
- 处理非结构化数据的挑战需要专家的人类监督才能有效实施.
相关概念视频
Steps in Outbreak Investigation
492
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
492
Statistical Methods for Analyzing Epidemiological Data
898
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
898
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K

